Two tracks, matching the two kinds of AI engineering interview. Track A covers how models are built, trained and served. Track B covers how products are built on top of them. Each page explains the material in moderate depth, puts a source link next to each claim so you can read further, and ends with a bank of interview questions with model answers.
Source chips like Author+ 2023 sit next to the claim they support. A number with a dashed underline is approximate, so check it against the primary source before quoting it. "May be out of date" boxes mark areas that change fast. The material was compiled in October 2026 from model knowledge plus web research, so for anything recent, check the current state.
Attention, MHA/MQA/GQA/MLA, positional encodings (RoPE, ALiBi), norms, activations, MoE, tokenization, long context, SSM/hybrids.
A2Data pipelines and filtering, objectives, scaling laws (Kaplan → Chinchilla → over-training), optimizers, LR schedules, stability, mid-training.
A3DP, ZeRO/FSDP, tensor, pipeline, sequence/context and expert parallelism, mixed precision (BF16/FP8), collectives, memory math, MFU.
A4SFT, reward models, RLHF with PPO, DPO and its variants, Constitutional AI / RLAIF, PEFT (LoRA/QLoRA), distillation, model merging.
A5RL fundamentals, policy gradients to PPO to GRPO, RLVR and reasoning models, building environments and verifiers, reward hacking, agentic RL infrastructure.
A6Prefill vs decode, KV cache, continuous batching, PagedAttention, prefix caching, speculative decoding, quantization, disaggregation, vLLM/SGLang/TRT-LLM.
A7GPU architecture, memory hierarchy, roofline, arithmetic intensity, kernel fusion, FlashAttention v1–3, Triton vs CUDA, tensor cores, GEMM tiling.
A8Benchmarks and their failure modes, contamination, pass@k, LLM-as-judge, test-time compute evals, and mechanistic interpretability basics (probes, SAEs, circuits).
Prompting, context engineering, structured outputs, tool calling, MCP, prompt caching, model routing, streaming, cost and latency levers.
B2Embeddings, chunking, ANN indexes (HNSW, IVF, PQ), vector DBs, BM25 and hybrid search, rerankers, query rewriting, GraphRAG, agentic RAG, retrieval evals.
B3Working, episodic, semantic and procedural memory; MemGPT/Letta, Mem0, Zep; extraction, consolidation, forgetting; context compaction; memory evals.
B4Workflows vs agents, ReAct, plan-and-execute, reflection, planner/orchestrator/worker, supervisor, handoffs, hierarchical and blackboard designs, frameworks, durable execution.
B5Eval-driven development, LLM-as-judge, tracing, guardrails, prompt injection, sandboxing, human-in-the-loop, reliability, cost control.
B6Worked interview designs: support agent, enterprise search / RAG, coding agent, deep-research agent, voice agent. Includes a framework for answering.
B7Claude Code, Codex, Cursor, Copilot and others: how they work, memory files, skills, subagents, hooks, MCP, daily workflows, team adoption, productivity evidence.
B8Always-on assistants like OpenClaw and Hermes: gateway, channel adapters, agent loop, memory, proactivity (heartbeats, cron), skills, and their security model.
B9Inside Claude Code, Codex, OpenCode, Pi and others: the agent loop, prompt assembly, tool and edit-format design, compaction, permissions and sandboxing, multi-model support, harness evals.
How to approach an LLD round, OOP and SOLID, the design patterns that come up, concurrency, worked classic problems (parking lot, LRU cache, rate limiter, elevator…) and AI-flavoured LLD problems.
C2Traditional building blocks (Kafka, CDC, queues, caches, sharding) combined with AI parts: streaming RAG ingestion, memory at scale, LLM gateways, agent platforms, LLM observability pipelines and more.
C3Web-scale data pipelines, GPU cluster scheduling, fault-tolerant training, RL rollout infra, multi-model inference platforms, eval and preference-data platforms, batch embedding backfills.