AI Engineer Interview Prep

Two tracks, matching the two kinds of AI engineering interview. Track A covers how models are built, trained and served. Track B covers how products are built on top of them. Each page explains the material in moderate depth, puts a source link next to each claim so you can read further, and ends with a bank of interview questions with model answers.

How to read these pages

Source chips like Author+ 2023 sit next to the claim they support. A number with a dashed underline is approximate, so check it against the primary source before quoting it. "May be out of date" boxes mark areas that change fast. The material was compiled in October 2026 from model knowledge plus web research, so for anything recent, check the current state.

Track A: Model internals (research / ML-systems interviews)

A1

Transformer & LLM Architecture

Attention, MHA/MQA/GQA/MLA, positional encodings (RoPE, ALiBi), norms, activations, MoE, tokenization, long context, SSM/hybrids.

A2

Pre-training

Data pipelines and filtering, objectives, scaling laws (Kaplan → Chinchilla → over-training), optimizers, LR schedules, stability, mid-training.

A3

Distributed Training & GPU Systems

DP, ZeRO/FSDP, tensor, pipeline, sequence/context and expert parallelism, mixed precision (BF16/FP8), collectives, memory math, MFU.

A4

Post-training & Alignment

SFT, reward models, RLHF with PPO, DPO and its variants, Constitutional AI / RLAIF, PEFT (LoRA/QLoRA), distillation, model merging.

A5

RL for LLMs & RL Environments

RL fundamentals, policy gradients to PPO to GRPO, RLVR and reasoning models, building environments and verifiers, reward hacking, agentic RL infrastructure.

A6

Inference & Serving

Prefill vs decode, KV cache, continuous batching, PagedAttention, prefix caching, speculative decoding, quantization, disaggregation, vLLM/SGLang/TRT-LLM.

A7

GPUs & Kernels

GPU architecture, memory hierarchy, roofline, arithmetic intensity, kernel fusion, FlashAttention v1–3, Triton vs CUDA, tensor cores, GEMM tiling.

A8

Model Evaluation & Interpretability

Benchmarks and their failure modes, contamination, pass@k, LLM-as-judge, test-time compute evals, and mechanistic interpretability basics (probes, SAEs, circuits).

Track B: AI product engineering (applied / agent interviews)

B1

LLM Application Foundations

Prompting, context engineering, structured outputs, tool calling, MCP, prompt caching, model routing, streaming, cost and latency levers.

B2

Semantic Search & RAG

Embeddings, chunking, ANN indexes (HNSW, IVF, PQ), vector DBs, BM25 and hybrid search, rerankers, query rewriting, GraphRAG, agentic RAG, retrieval evals.

B3

Agent Memory Systems

Working, episodic, semantic and procedural memory; MemGPT/Letta, Mem0, Zep; extraction, consolidation, forgetting; context compaction; memory evals.

B4

Agent Architectures

Workflows vs agents, ReAct, plan-and-execute, reflection, planner/orchestrator/worker, supervisor, handoffs, hierarchical and blackboard designs, frameworks, durable execution.

B5

Evals, Observability & Safety in Production

Eval-driven development, LLM-as-judge, tracing, guardrails, prompt injection, sandboxing, human-in-the-loop, reliability, cost control.

B6

AI System Design Case Studies

Worked interview designs: support agent, enterprise search / RAG, coding agent, deep-research agent, voice agent. Includes a framework for answering.

B7

Working with AI Coding Tools

Claude Code, Codex, Cursor, Copilot and others: how they work, memory files, skills, subagents, hooks, MCP, daily workflows, team adoption, productivity evidence.

B8

Personal AI Assistants

Always-on assistants like OpenClaw and Hermes: gateway, channel adapters, agent loop, memory, proactivity (heartbeats, cron), skills, and their security model.

B9

Harness Engineering

Inside Claude Code, Codex, OpenCode, Pi and others: the agent loop, prompt assembly, tool and edit-format design, compaction, permissions and sandboxing, multi-model support, harness evals.

Track C: Design interviews (LLD and system design)

C1

Low-Level Design (LLD)

How to approach an LLD round, OOP and SOLID, the design patterns that come up, concurrency, worked classic problems (parking lot, LRU cache, rate limiter, elevator…) and AI-flavoured LLD problems.

C2

System Design: AI Products on Real Infrastructure

Traditional building blocks (Kafka, CDC, queues, caches, sharding) combined with AI parts: streaming RAG ingestion, memory at scale, LLM gateways, agent platforms, LLM observability pipelines and more.

C3

System Design: Training & ML Infrastructure

Web-scale data pipelines, GPU cluster scheduling, fault-tolerant training, RL rollout infra, multi-model inference platforms, eval and preference-data platforms, batch embedding backfills.

Suggested study order