Track B · AI product engineering

AI System Design Case Studies

The "design an AI system" round is where everything else in this guide gets tested together: retrieval, agents, evals, serving, cost and safety. Interviewers aren't looking for a perfect diagram. They want to see that you can turn a vague product ask into a measurable task, start from the simplest thing that could work, do the token and dollar math out loud, and design for a component that is non-deterministic and sometimes wrong. This page gives you a reusable framework, then works through six full case studies and three short ones in the shape a strong candidate would present them.

TL;DR: the 8–12 things to be able to say out loud

  • Clarify the task as an eval before drawing boxes. "What does a correct output look like, who judges it, and what's the cost of a wrong one?" Everything downstream depends on this.
  • Start with a one-prompt baseline, then add retrieval, tools, multiple steps or multiple agents only when a measured failure justifies it.
  • The quality / latency / cost triangle is the core trade-off. Every lever (bigger model, more context, more steps, more samples) buys quality with latency and money.
  • Cost per request = tokens × price, and agent loops multiply tokens because the context is re-sent every step. Prompt caching (cache reads at roughly a tenth of the input price) and model routing are the first two cost levers.
  • Agents re-read their whole context every turn, so input tokens dominate cost; output tokens dominate latency.
  • Permissions and side effects live in deterministic code, not in the prompt. Tools enforce auth, limits and idempotency; the model only proposes.
  • Evals are the spec: offline golden sets plus LLM-as-judge with calibrated rubrics, outcome-based checks (database state, tests passing), and online metrics (resolution, escalation, thumbs, cost).
  • Prompt injection is the default threat model for any system that reads untrusted content and can act. Break the "lethal trifecta": private data + untrusted input + an exfiltration channel.
  • Human-in-the-loop is a design component: confidence thresholds, approval gates for irreversible actions, and graceful hand-off with full context.
  • Voice is a latency problem: about 1–1.5 s voice-to-voice is a realistic target, and turn detection matters as much as model speed.
  • Memory is a data-governance problem as much as a retrieval problem: what to store, how to update or forget, and how the user sees and edits it.
  • Close with an iteration plan: ship to a slice, log traces, mine failures into the eval set, and tune the cheapest lever first.

A reusable framework for AI system design interviews

Classic system design interviews follow a familiar arc: requirements, API, data model, high-level design, deep dives, scaling. AI system design keeps that arc but adds steps for the parts that are new: defining quality, choosing and composing models, evaluating a stochastic component, and paying per token. Here is a ten-step framework you can run through in roughly 40 minutes. You won't spend equal time on each step. The interviewer will usually pull you into two or three deep dives, so keep the rest brief and signpost what you're skipping.

#StepWhat you actually say or produceTime
1Clarify the task and success metricsWho the users are, the inputs and outputs, what "correct" means, and the business metric (resolution rate, time saved, conversion). Turn that into an eval definition.4–6 min
2ConstraintsLatency (p50/p95, streaming or not), cost ceiling per request or per month, scale (DAU, QPS, corpus size), privacy and residency, the accuracy bar and the cost of errors, regulatory needs.2–3 min
3BaselineThe simplest thing that could work: one prompt to a capable model with the obvious context pasted in. Say what it would get right and where it would fail.2 min
4Data and context sourcesWhat the model needs to know (docs, DB rows, user history, tools), freshness, access control, volume. Retrieval vs tools vs fine-tuning.3–5 min
5ArchitectureDiagram: entry point, orchestrator, retrieval, tools, model calls, state store, guardrails, human hand-off, logging. Workflow (fixed steps) vs agent (model picks the steps).8–10 min
6Model choicesWhich tier per step (small/fast for routing and classification, frontier for hard reasoning), hosted API vs open weights, fine-tune or not, structured outputs.2–3 min
7EvalsOffline golden set, LLM judge plus rubric, outcome checks, regression gates in CI, online metrics and A/B, human review sampling.4–6 min
8Failure modes and safetyHallucination, wrong tool call, prompt injection, data leakage, runaway loops, provider outage. A mitigation for each one.3–5 min
9Scaling and cost mathQPS, tokens per request, dollars per day, GPU or rate-limit needs, caching, batching, routing.3–5 min
10Iteration planRollout (shadow, canary, % ramp), the feedback loop from traces to the eval set, what you'd build in v2.2 min
Clarifytask+metric Constr.lat/$/scale Baseline1 prompt Contextdata/tools Arch.diagram Modelstiers Evalsthe spec Failures+safety CostQPS, $/day Iterateship+learn production failures become new eval cases evals can reveal the task was mis-specified
The ten-step framework. It's a loop, not a waterfall: the eval definition from step 1 is reused in step 7, and production traces feed it in step 10.

Step 1 in more depth: turn the ask into an eval

The most common weak opening is jumping to "we'll use RAG with a vector DB." The strong opening is a set of questions that pin down the task as something you could grade:

Then say it as a sentence: "We'll call v1 successful if, on a golden set of 500 real tickets, it fully resolves at least 60% with no policy violations, escalates the rest with a correct summary, keeps p95 first-token latency under 2 s, and costs under $0.30 per conversation." The numbers are placeholders you agree with the interviewer, but having them changes the whole conversation.

Step 3 in more depth: why the baseline matters

Anthropic's guidance on agents is to find the simplest solution and add complexity only when it demonstrably helps. Their taxonomy separates workflows (LLMs and tools orchestrated through predefined code paths: prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer) from agents (the model dynamically directs its own process and tool use). Anthropic 2024 Interviewers reward candidates who climb that ladder explicitly:

one prompt → + retrieval / pasted context → + tools (single call) → fixed workflow (chain / route) → agent loop (model chooses tools) → multi-agent (orchestrator + workers) Each rung adds: capability ↑ latency ↑ cost ↑ variance ↑ eval difficulty ↑ Climb only when an eval shows the lower rung failing for a reason the higher rung fixes.

A related data point from coding: the "Agentless" paper showed that a fixed three-phase pipeline (localize, repair, validate) beat the open-source agents of the time on SWE-bench Lite, at about $0.70 per issue. Xia+ 2024 Agents have improved a lot since then, but the lesson holds: a well-structured workflow is often a strong baseline, and you should be able to argue for it.

How AI system design differs from classic system design

DimensionClassic system designAI system design
CorrectnessDeterministic. A request either succeeds or errors.Stochastic and graded. The same input can give different outputs, and "200 OK" can still be wrong. Quality is a distribution you measure.
SpecAPI contract and requirements doc.The eval set is the spec. You can't say a change is an improvement without running it.
Unit costFractions of a cent per request; compute is amortized.Cents to dollars per request, scaling linearly with tokens. Cost is a first-class design constraint, often the binding one.
LatencyMilliseconds, dominated by I/O and DB.Seconds, dominated by token generation (roughly proportional to output length) and sequential model calls.
Main trade-offConsistency / availability / latency.Quality / latency / cost, plus safety. Bigger models, more context and more steps buy quality at the price of the other two.
Failure modesCrashes, timeouts, data races.Confident wrong answers, wrong tool arguments, prompt injection, runaway loops, silent quality regressions when a provider updates a model.
TestingUnit and integration tests with exact asserts.Statistical evals, LLM judges, pass@k and pass^k, human review sampling, online A/B.
DependenciesYour own services.Often an external model API with rate limits, outages, deprecations and price changes. A gateway and fallbacks are needed.
SecurityAuthn/z, input validation.All of that, plus: the model reads attacker-controlled text and may follow it. Treat model output as untrusted input to your tools.
Quality Latency Cost bigger model · more context · more steps reasoning tokens · self-consistency · verifier buy quality, spend latency and $ caching · routing · distillation · batching streaming · parallel calls · shorter outputs claw back latency and $, at some risk to quality
The quality / latency / cost triangle. Name the lever and which corner it pulls toward.

What interviewers actually grade

Signals of a strong candidate

  • Asks clarifying questions that change the design (error cost, scale, latency mode).
  • States success metrics and an eval plan early.
  • Proposes a simple baseline and justifies each added component.
  • Does token and dollar math with stated assumptions, out loud.
  • Puts auth, limits and irreversible actions in code, not prompts.
  • Names specific failure modes with specific mitigations.
  • Knows rough magnitudes: TTFT, tokens/s, price tiers, embedding sizes.
  • Has an iteration and rollout plan; treats production as the data source.

Signals of a weak one

  • Starts with "multi-agent system with a vector DB" before knowing the task.
  • "We'll fine-tune" as a reflex, with no data plan.
  • No numbers: no QPS, tokens or dollars.
  • "The LLM will check permissions."
  • Evals amount to "we'll test it."
  • Ignores latency for user-facing paths, or cost for agents.
  • No human fallback; no plan for when the model is wrong.
  • Buzzword lists without trade-offs.

Back-of-envelope toolkit

You'll reuse the same handful of formulas in every case study. Memorize the shape, not the exact prices, which change every few months.

$$ \text{cost/request} = \sum_{\text{calls}} \Big( T_{\text{in,uncached}}\,p_{\text{in}} + T_{\text{in,cached}}\,p_{\text{cache}} + T_{\text{out}}\,p_{\text{out}} \Big) $$ $$ \text{QPS}_{\text{avg}} = \frac{\text{requests/day}}{86{,}400}, \qquad \text{QPS}_{\text{peak}} \approx (2\text{–}5)\times \text{QPS}_{\text{avg}} $$ $$ \text{latency} \approx \text{TTFT} + \frac{T_{\text{out}}}{\text{tokens/s}} \quad\text{(per sequential call)}, \qquad \text{concurrency} = \text{arrival rate} \times \text{duration} $$

The last one is Little's law, and it's what sizes long-running agents and voice calls: 100 calls/s arriving that each last 4 minutes means about 24,000 concurrent calls.

Rule of thumbValue to use (state it as an assumption)
Tokens per English wordAbout 1.3 (roughly 4 characters per token). Newer tokenizers vary; Anthropic notes its Claude 4.7+ tokenizer produces about 30% more tokens for the same text. Anthropic pricing
Mid-tier hosted model list price (late 2026)$2–3 per M input tokens, $10–15 per M output. On this page I use $3 / $15 as a round, slightly conservative number.
Small / fast modelAbout $1 / $5 per M (e.g. Haiku-class) Anthropic pricing
Frontier modelAbout $4–5 / $20–25 per M, with premium tiers above that
Prompt-cache readAbout 0.1× the base input price on Anthropic's API (5-minute cache write costs 1.25×). Anthropic pricing Other providers have similar discounts with different mechanics.
Batch API discount50% for asynchronous jobs on Anthropic's API Anthropic pricing
Time to first token (hosted, short prompt)200–800 ms; grows with prompt length and reasoning
Output speed50–150 tokens/s for typical hosted models; specialized inference hardware can be much faster
Embedding vector1024 dims × 4 bytes = 4 KB float32; about 1 KB int8; 128 bytes binary
May be out of date

Model prices dropped steeply through 2024–2026 and model names turn over every few months. The pricing figures above come from Anthropic's pricing page as of October 2026. In an interview, say "assuming roughly $3/$15 per million tokens" and move on. The interviewer cares about the structure of the calculation, not the third significant figure.

Interview angle

A common probe is "how would you cut the cost by 10×?" A strong answer ranks levers by effort and risk: (1) prompt caching for the static prefix, which is nearly free; (2) trimming context with better retrieval and shorter tool outputs; (3) routing easy traffic to a small model, measured by evals; (4) batching non-interactive work at 50% off; (5) distilling or fine-tuning a small model on logged traffic once you have volume. Then say what you'd measure to make sure quality held.

Go deeper

Case study 1: Customer-support agent with tool access and human escalation

Prompt: "Design an AI agent for an e-commerce company that handles customer support chat. It should be able to look up orders, process refunds and cancellations, and escalate to human agents when needed."

This is the most common AI design question because it touches everything: conversation, retrieval over policy docs, tool calling with side effects, authorization, escalation, and a clear business metric. Anthropic uses customer support as the canonical example of where agents fit, because it pairs conversation with actions and success is measurable by resolution. Anthropic 2024

Requirements and clarifying questions

QuestionWhy it mattersAssumed answer
Channels: web chat only, or email and voice too?Latency budget and turn structure differ a lot.Web and in-app chat first; email later.
Is the user authenticated?Determines whether the agent can act at all, and how identity is bound to tools.Yes, logged-in users. Guests get FAQ only.
Which actions, and with what limits?Side effects need policy limits and approval gates.Order lookup, tracking, cancel before shipment, refunds up to $200 automatically; above that needs human approval.
Volume and peaks?Sizing, rate limits, cost.50k conversations/day, 3× peak on sales days.
Current baseline?Defines success.Human agents; average cost per contact $4–6; CSAT 4.1/5.
Languages and regions?Model choice, policy variants, data residency.English and Spanish; US and EU (GDPR).
What does the hand-off look like?Integration with the ticketing tool, agent desktop.Escalate into the existing helpdesk queue with summary and transcript.

Success metrics. Primary: automated resolution rate (conversation ends with the issue solved and no human touch, and no re-contact within 7 days). Guardrails: CSAT not below baseline, policy-violation rate (refunds outside policy) below 0.5%, wrong-action rate near zero, escalation quality (the human didn't have to re-ask). Operational: p95 time-to-first-token under 2 s, cost per conversation under $0.40.

Back-of-envelope

Assumptions (state them explicitly): 50k conversations/day; 6 user turns per conversation; an average of 2 model calls per user turn (one to decide a tool call, one to answer after the tool returns), so 12 calls per conversation; a mid-tier model at $3 / $15 per M tokens.

QuantityCalculationResult
Model calls/day50k × 12600k
Average QPS (model calls)600k / 86,400about 7
Peak QPS× 3 on sale daysabout 20
Input tokens per callsystem prompt + policies 3k, tool schemas 2k, conversation so far ~2k avg, retrieved policy snippets 1.5k~8.5k
Output tokens per calltool-call JSON or a short reply~250
Per conversation, uncached12 × 8.5k = 102k in → $0.31; 12 × 250 = 3k out → $0.045≈ $0.35
Per conversation, with prompt cachingthe 5k static prefix is cached on 11 of 12 calls: 55k tokens at $0.30/M ≈ $0.017; the rest (~47k) at $3/M ≈ $0.14; output $0.045≈ $0.20
Daily model spend50k × $0.20–0.35≈ $10k–17.5k/day
Routing the easy ~40% (FAQ, tracking) to a small model at ~1/3 the priceanother ~25% saving

Compare to the human baseline: if 60% of 50k conversations are automated at roughly $0.25 each, that replaces about 30k human contacts a day. At $4–6 each, that's well over $100k/day of human cost against roughly $12k of model spend. The economics are rarely the blocker here; trust, correctness and policy compliance are. Say that out loud: it tells the interviewer where you'll spend design effort.

Architecture

Chat widgetauthn session Gatewayauth · rate limit Input guardPII, abuse, injection Orchestratorintent router → agent loopstate machine + budget LLM (via LLM gateway)small router + mid agent Policy / FAQ RAGhybrid search, versioned Tool serviceuser-scoped tokenspolicy engine, limitsidempotency keys Orders API Payments/refunds Approval queuerefund > limit Escalationsummary → helpdesk Human agentsagent desktop Conversation store + traces + eval loggingevery prompt, tool call, result, decision
The model proposes, the tool service disposes. Identity, limits and approvals are enforced outside the model.

Component walkthrough

Key design decisions and alternatives

DecisionChosenAlternativeWhy
Control flowHybrid: router + fixed workflows for the top intents, agent loop for the long tailPure agent for everythingFixed flows are cheaper, faster and easier to eval for the 60–70% of traffic that's repetitive. The agent handles the messy remainder.
Where policy livesDeterministic policy engine in the tool layer, plus policy text in context for explanationPolicy only in the promptPrompts are advisory; code is enforcement. The Air Canada tribunal case is the standard cautionary tale: the airline was held liable when its chatbot told a customer a bereavement refund could be claimed retroactively, contrary to its actual policy. Dentons 2024
Single vs multi-agentSingle agent with ~10 toolsSpecialist agents (billing, shipping) with a supervisorMulti-agent adds hand-off failures and latency. Split only if the tool count or prompt size grows enough to hurt tool selection accuracy.
ModelSmall model for routing and guards; mid-tier for the agent loopFrontier everywhereTool-use reliability matters more than raw reasoning here; measure on the eval set before upgrading.
Fine-tuningNot in v1Fine-tune on historic transcriptsHistoric human transcripts encode old policies and human shortcuts. Revisit for routing and tone once there's logged, labeled agent traffic.
Common mistake

Letting the model pass customer_id as a tool argument. A prompt-injected or confused model can then read someone else's orders. Bind identity on the server; tools take only the resource ID, and the tool layer checks that the resource belongs to the session's customer.

Evals

Failure modes

FailureMitigation
Hallucinated policy ("you can return it after 90 days")Ground answers in retrieved policy and cite the section; deterministic eligibility tool; judge-based spot checks for unsupported claims.
Wrong action (refunds the wrong item)Confirmation step before mutations ("I'll refund the blue jacket, $59. Shall I proceed?"); idempotency keys; limits; reversible actions preferred.
Prompt injection via order notes or product reviewsTreat tool outputs as data; the tool layer enforces authorization regardless of what the model asks; no tool that sends data to arbitrary destinations.
Loops (agent keeps calling the same tool)Step and token budgets; detect repeated identical calls; escalate on budget exhaustion.
Provider outage or latency spikeLLM gateway with fallback model; degrade to "FAQ plus escalate" mode.
Abuse and fraud (refund farming)Fraud score in the eligibility tool; velocity limits per account; human approval above thresholds.

How to extend

Interview angle

The deep dive is almost always "how do you stop it doing something bad?" Strong answers layer the defenses: identity bound server-side; narrow tools; a deterministic policy engine; confirmation before mutations; amount limits and human approval; idempotency; full audit logs; and evals that check database outcomes, not just text. Mention that the company is legally responsible for what the bot says, so grounding and conservative escalation are product requirements, not nice-to-haves.

Go deeper

Case study 2: Enterprise knowledge search (RAG) over Confluence, Drive and Slack with permissions

Prompt: "Employees waste hours finding information spread across Confluence, Google Drive and Slack. Design an assistant that answers questions over all of it, with citations, and never shows someone a document they can't access."

The interesting part isn't the LLM call. It's the ingestion pipeline, permission enforcement, freshness, and retrieval quality on messy enterprise data. Treat "never leaks a document" as a hard constraint, on par with correctness.

Requirements and clarifying questions

Back-of-envelope

Assumptions: 20k users × 10 queries/day; chunks of about 500 tokens; 1024-dim embeddings; answer generation with a mid-tier model at $3 / $15 per M tokens.

QuantityCalculationResult
Queries/day20k × 10200k
QPS200k / 86,400, peak ×4 in business hoursabout 2.3 avg, ~10 peak
Chunks20B tokens / 50040M chunks
Vector storage40M × 4 KB (float32) / 1 KB (int8) / 128 B (binary)160 GB / 40 GB / 5 GB
Initial embedding cost20B tokens × $0.02–0.10 per M$400–2,000 one-time
Optional contextual chunk headersAnthropic estimated $1.02 per M document tokens with prompt caching Anthropic 2024; × 20B≈ $20k one-time (prices have fallen since, so likely less now)
Tokens per answer10 chunks × 500 + instructions and history ~2k → ~7–8k in; ~400 out
Cost per answer8k × $3/M + 400 × $15/M≈ $0.03
Daily generation cost200k × $0.03≈ $6k/day (≈ $2k with a small model for easy queries)
Incremental ingestionassume 1% of the corpus changes per day: 200M tokenssmall: tens of dollars of embeddings per day

Takeaway: generation dominates cost and retrieval quality dominates correctness; storage is cheap. 40M vectors fits comfortably in a single sharded vector index, or in Postgres with pgvector at the low end, and comfortably in a dedicated vector DB or a search engine with vector support (OpenSearch, Elasticsearch, Vespa).

Architecture

INGESTION (async, continuous) Confluence Google Drive Slack Connectorswebhooks + pollingcontent + ACLs + ids Parse + chunkstructure-aware, threads Identity syncusers, groups → principals Embed + enrichcontext header, metadata Hybrid indexBM25 + vectorsacl_principals[] per chunksource, date, author QUERY (online, per request) User (SSO)principal set Query plannerrewrite, filters, route Retrieve top-100ACL filter IN the query Rerank → top-10cross-encoder Live ACL re-checksource API, cached 5 min LLM answer + citesstreamed, grounded Citation verifyclaim ↔ chunk
Two pipelines. Permissions are enforced at retrieval time (pre-filter on principals) and re-checked against the source for the final few documents.

Component walkthrough

Key design decisions and alternatives

DecisionOptionsRecommendation and reasoning
Permission enforcement point(a) Post-filter results after retrieval; (b) pre-filter inside the index query; (c) separate index per user or group(b) plus a live re-check. Post-filtering can return zero results when the top-k is mostly forbidden, and leaks via timing or counts. Per-user indexes don't scale with 20k users and overlapping groups.
RetrievalDense only; BM25 only; hybridHybrid plus a reranker. Enterprise queries are full of rare identifiers.
Agentic vs single-shot retrievalOne retrieval call vs an agent that searches iterativelySingle-shot for most questions (fast, cheap). Escalate to an agentic multi-search mode for complex "compare / summarize across" questions, with a visible "researching…" state.
Long context instead of RAGStuff whole documents into a 1M-token windowUseful for "summarize this doc" on one or a few docs. Not a substitute at 20B tokens, and models use the middle of long contexts less reliably. Liu+ 2023
Fine-tune on the corpusTeach the model company factsNo. Facts change daily, fine-tuning can't enforce per-user permissions, and it can't cite. Retrieval is the right tool for knowledge; fine-tuning is for behavior and format.
Slack handlingIndex each message; index threads; index channel-day windowsThreads as units, with channel name and date. Single messages lack context; day windows mix topics.
Common mistake

Embedding documents with a service account and enforcing permissions "in the prompt" ("only use documents the user can see"). The model has no idea what the user can see, and anything in the context window can end up in the answer. Permissions must be enforced before content reaches the model.

Evals

Failure modes

FailureMitigation
Stale or conflicting documents (three versions of the travel policy)Boost recency and "official" spaces; dedupe near-duplicates; show dates in citations; tell the model to surface conflicts rather than pick silently.
Permission leak through ACL lagEvent-driven revocation sync; live re-check on final documents; short cache TTLs.
Over-shared content (a doc shared "anyone in the company" that shouldn't be)Not technically a leak, but AI search makes it findable. Offer admins an over-sharing report and sensitive-label exclusions before launch.
Prompt injection in an indexed doc ("ignore previous instructions and…")Treat retrieved text as data, delimit it, give the answering model no side-effecting tools, and sanitize rendered links and images (no auto-loading images that could exfiltrate data through URLs).
Hallucinated answer when retrieval missesRelevance threshold on reranker scores; abstain below it; citation verification.
Connector rate limits and API changesBackoff, per-source queues, health dashboards, and alerting on ingestion lag.

How to extend

Interview angle

Expect a deep dive on permissions: "a user loses access to a doc; walk me through what happens." A strong answer covers the revocation event from the source, the identity and ACL sync, updating the chunk's principals in the index, the live re-check as a safety net, and what the worst-case window is. Also expect "how do you know retrieval is the problem and not generation?", which you answer with separate retrieval metrics and by looking at whether the right chunk was in the context for the failed answers.

Go deeper

Case study 3: Autonomous coding agent (issue → pull request)

Prompt: "Design a system where engineers assign a GitHub issue to an AI agent, and it comes back with a pull request that passes CI."

Coding is the domain where agents work best today, for one structural reason: there's a cheap, automatic verifier. Tests, type checkers, linters and builds give the agent a feedback signal it can iterate against, and give you an eval. Anthropic's coding-agent guidance boils down to "give the agent a check it can run," because without one, "looks done" is the only stopping signal. Claude Code docs

Requirements and clarifying questions

Back-of-envelope

Assumptions: a typical task takes about 50 agent steps (read files, grep, edit, run tests); the context grows from ~10k to ~80k tokens, averaging ~40k input tokens per step; ~400 output tokens per step; a mid-tier model at $3 / $15 per M; prompt caching on the growing prefix.

QuantityCalculationResult
Input tokens per task50 steps × 40k~2M
Output tokens per task50 × 400~20k
Uncached cost per task2M × $3/M + 20k × $15/M≈ $6.30
Cached cost per task~90% of input is a cache hit (the prefix from previous steps): 1.8M × $0.30/M ≈ $0.54; ~0.2M new tokens written at 1.25× ≈ $0.75; output $0.30≈ $1.60
Daily model spend2,000 × $1.60–6.30≈ $3k–13k/day
Wall-clock per task50 steps × (~3 s model + ~3 s tool/test)~5 min model time, often 10–30 min with test suites
Concurrent sandboxes2,000/day × 20 min ÷ 1,440 min/day, peak ×3~28 avg, ~100 peak
Sandbox compute2 vCPU / 4 GB for 20 min at ~$0.05–0.10 per vCPU-houra few cents per task: negligible next to tokens

Two important observations. First, caching is the difference between viable and not: agent loops re-send the whole growing transcript every step, so the uncached cost grows roughly quadratically with step count. Second, the right denominator is cost per merged PR, not cost per task: if only 40% of PRs merge, the effective cost is 2.5× higher, and a stronger, pricier model with a higher merge rate can be cheaper overall.

$$ \text{cost per merged PR} = \frac{\text{cost per attempt}}{\text{merge rate}} + \text{reviewer time} \times \text{engineer hourly cost} $$

Reviewer time usually dominates: 15 minutes of an engineer's time is worth far more than the tokens. So optimize for PRs that are easy to review (small, well-described, with tests), not just for raw success rate.

Architecture

GitHub issue / Slack / CLI ──► Task API ──► Task queue (durable, per-repo concurrency limits) │ ▼ ┌─────────────────────────────┐ │ Orchestrator (durable job) │ budget: steps, tokens, wall-clock │ plan → act → verify loop │ checkpoints after each step └──────┬───────────────┬──────┘ model calls │ │ tool calls (structured) ┌─────────────────────▼──┐ ┌────────▼──────────────────────────────┐ │ LLM gateway │ │ Sandbox (microVM / gVisor container) │ │ mid/frontier for agent │ │ repo checkout @ base SHA (worktree) │ │ small model: summarize │ │ tools: read/grep/glob, edit (diff), │ │ prompt caching on │ │ bash (allowlisted), run_tests │ └─────────────────────────┘ │ network: egress allowlist (pkg mirrors)│ │ NO prod creds, scoped GitHub token │ └────────┬──────────────────────────────┘ │ diff + test results + transcript ▼ Verifier stage: full test suite, lint, typecheck, (optional) reviewer-agent in fresh context │ pass │ fail after N tries ▼ ▼ Open draft PR (summary, test evidence) Report back with findings │ ▼ Human review ──► merge (human only) Repo context: code index (symbols, embeddings), AGENTS.md / CLAUDE.md conventions, past PRs

Component walkthrough

Key design decisions and alternatives

DecisionOptionsRecommendation and reasoning
Agent vs fixed pipelineAgentless-style localize → repair → validate vs open-ended agentAgent with a suggested plan structure. The fixed pipeline is a strong, cheap baseline (Agentless: 32% on SWE-bench Lite at $0.70/issue in 2024 Xia+ 2024), but real repos need running builds, reading errors and adapting.
Sampling strategyOne attempt vs N parallel attempts with selectionStart with one. For hard tasks, run 2–4 attempts in parallel sandboxes and pick the one that passes tests with the smallest diff. This trades cost for success rate, and test-based selection is reliable.
IsolationDocker container; gVisor; microVMMicroVM or gVisor for untrusted, model-generated code. Plain containers share the host kernel.
NetworkOpen; allowlist; noneAllowlist (package mirrors, internal docs). Open network plus repo secrets plus untrusted issue text is the lethal trifecta. Willison 2025
CredentialsDeveloper's token vs scoped app tokenA GitHub App token scoped to one repo, able to push to agent branches only. No deploy or production credentials in the sandbox at all.
ModelFrontier vs mid-tierMeasure cost per merged PR. Coding is where frontier models most often pay for themselves.

Evals

May be out of date

Top SWE-bench Verified scores rose from the low teens (SWE-agent reported 12.5% on full SWE-bench in 2024 Yang+ 2024) to well above 70% by 2025–2026, and the community has raised concerns about saturation and contamination of the public set. Check the current leaderboard and newer benchmarks before quoting numbers, and in interviews, emphasize internal evals over public ones.

Failure modes

FailureMitigation
Gaming the tests (deleting failing tests, special-casing inputs, weakening asserts)Verifier flags test-file deletions and modified asserts; reviewer agent instructed to look for this; protected test directories; hidden held-out tests in evals.
Wandering (reads 200 files, never converges)Step budget; plan-first prompting; sub-agents for exploration; stop and report when the budget runs out.
Context overflow on large reposCompaction; truncated tool outputs; just-in-time file reads rather than pre-loading.
Prompt injection through issue text, code comments or dependency READMEsNo secrets in the sandbox; egress allowlist; scoped tokens; human merge gate.
Flaky tests misleading the agentRun the baseline suite before changes to record pre-existing failures; rerun failures once.
Huge, unreviewable PRsDiff-size limit; ask the agent to split; decline tasks that triage scores as too big.

How to extend

Interview angle

Interviewers probe the sandbox ("what stops it from curl-ing your secrets to an attacker?") and the verification story ("how do you know the PR is correct?"). A strong answer separates the agent's own checks from an independent verifier, mentions test gaming as a known failure, and frames the metric as cost per merged PR including reviewer time. Bonus: explain why caching makes or breaks the economics of long agent loops.

Go deeper

Case study 4: Deep-research agent (orchestrator with parallel web-search subagents and citations)

Prompt: "Design a 'deep research' feature: the user asks a complex question, and the system spends several minutes searching the web and produces a long, well-cited report."

This is the canonical case for multi-agent systems, and there's an unusually detailed public write-up: Anthropic's description of how they built Claude's Research feature. Anthropic 2025 OpenAI's deep research, launched in early 2025, was described as an agent powered by a version of o3 optimized for browsing, searching and analyzing large amounts of web content. OpenAI 2025 Use these as reference points, but design from first principles.

Requirements and clarifying questions

Back-of-envelope

Assumptions: a lead agent plus 5 subagents on average; each subagent makes about 20 tool calls (searches and page fetches) and averages about 30k tokens of context per step; a final citation pass; mid-tier model at $3 / $15 per M; web search at $10 per 1,000 searches (Anthropic's listed price for its server-side search tool). Anthropic pricing

ComponentInput tokensOutput tokens
Lead agent: plan, dispatch, read results, synthesize~200k~15k (the report)
5 subagents × 20 steps × ~30k context~3M5 × 10k = 50k
Citation pass~100k~5k
Total~3.3M~70k

Architecture

User query+ clarify Qs Lead agentplan → subtasksevaluate gaps → re-dispatchsynthesize report Memory / scratchpadplan, findings, sources Subagent 1"pricing of X, Y, Z" Subagent 2"benchmarks" Subagent 3"customer reports" Subagent Nown context window Web search API Fetch + extractHTML/PDF → text Source store Code exectables, math Citation agentmap claims → exact sources Report + citationsstreamed progress to UI condensed findings
Orchestrator-worker. Subagents explore in parallel with their own context windows and return compressed findings; a separate pass attaches citations.

Component walkthrough

Key design decisions and alternatives

DecisionOptionsRecommendation and reasoning
Single agent vs multi-agentOne long-context agent doing everything sequentially vs orchestrator plus parallel workersMulti-agent fits here because research is breadth-first and parallelizable with mostly independent subtasks. Anthropic's multi-agent setup (Opus 4 lead, Sonnet 4 subagents) beat single-agent Opus 4 by 90.2% on their internal research eval, and token usage alone explained 80% of the variance on BrowseComp. Anthropic 2025
When not to use multi-agentTasks where steps depend tightly on each other, or where all agents need shared context, such as most coding. Cognition argued that parallel subagents make conflicting implicit decisions unless they share full context and traces. Yan 2025 The two posts aren't really contradictory: they describe different task shapes.
Model mixSame model everywhere vs strong lead plus cheaper workersA strong lead (planning and synthesis are the hard parts) with mid-tier workers. Measure whether upgrading workers moves the eval.
Search providerCommercial search API, provider-native search tool, own crawlerCommercial or provider-native search for v1. Your own index only for a vertical (legal, academic) where you control the corpus.
Report generationSingle pass vs outline then section-by-sectionOutline first, then sections, for long reports: it keeps each generation call focused and lets you stream progress.
Intuition

Multi-agent systems are mostly a way to spend more tokens productively in parallel. If the task's quality scales with how much material you read and the subtasks are separable, more parallel context windows win. If the task needs one coherent line of reasoning, they mostly add coordination failures.

Evals

Failure modes

FailureMitigation
Fabricated or mismatched citationsSeparate citation pass; only allow citing fetched sources from the source store; automated quote-presence checks.
Low-quality sources (SEO spam, AI-generated content farms)Source-quality heuristics in the subagent prompt; domain allow/deny lists; prefer primary sources.
Over-spawning (50 subagents for a simple question)Explicit effort-scaling rules in the lead prompt; hard caps on subagents and tool calls; budget per report.
Duplicated or overlapping subagent workClear task boundaries in delegation; lead dedupes findings.
Prompt injection from web pagesWeb content is untrusted; subagents have no side-effecting tools; never combine web browsing with access to private user data and an outbound channel without strict controls.
Errors compounding over a long runCheckpoints and resume; let the agent know when a tool failed so it can adapt; tracing of every decision for debugging.

How to extend

Interview angle

Expect "why multi-agent instead of one agent with a big context?" The strong answer: context isolation (each worker can read 100k+ tokens and return a summary), parallelism (wall-clock time), and the evidence that more tokens spent productively correlates with quality on this task type. Then volunteer the costs: roughly 15× chat tokens, coordination failures, harder debugging, and that it's a poor fit for tightly coupled tasks. Showing both sides is the signal.

Go deeper

Case study 5: Real-time voice agent

Prompt: "Design a voice agent that answers inbound phone calls for a chain of clinics: it books, moves and cancels appointments, and answers basic questions."

Everything from case study 1 applies (tools, policy, escalation), but voice adds a hard real-time constraint. Humans expect a reply within a few hundred milliseconds of finishing a sentence; a 3-second pause feels broken. The design is dominated by the latency budget, turn-taking, and telephony plumbing.

Requirements and clarifying questions

Two architectures: cascaded pipeline vs speech-to-speech

Cascaded: STT → LLM → TTS

  • Streaming speech-to-text produces partial and final transcripts.
  • A text LLM with tools generates a streamed reply.
  • Streaming text-to-speech starts speaking from the first sentence fragment.
  • Pros: pick the best component for each stage; full text transcript for compliance, evals and debugging; mature tool calling; easy context control; cheaper.
  • Cons: more moving parts; latency adds up across stages; loses prosody (tone, hesitation) from the input.

Speech-to-speech (native audio model)

  • One model takes audio in and produces audio out (e.g. OpenAI's Realtime API over WebRTC or WebSocket). OpenAI docs
  • Pros: potentially lower latency; more natural prosody and emotion; hears tone and non-verbal cues.
  • Cons: one practitioner guide estimates it at 3–5× the cost of a cascaded pipeline, with less flexible context management and no reliable transcript for compliance Voice AI primer 2026; harder to eval; tool-use reliability historically behind text models.

For a clinic-booking agent with compliance needs and heavy tool use, the cascaded pipeline is the defensible default in 2026. Speech-to-speech fits consumer companionship or language-practice products where naturalness beats auditability. Say that this is a fast-moving trade-off.

May be out of date

Speech-to-speech models improved rapidly in 2025–2026, and the cost and tool-calling gap has been narrowing. Hybrid designs (audio-native understanding with text-based reasoning and tools) are also appearing. Check current model capabilities and pricing before committing to either side in a real design.

Latency budget

The most useful public breakdown comes from a widely used practitioner primer (originally written for the AI Engineering Summit in February 2025, updated in 2026), which totals roughly 1.3 s for a typical cascaded pipeline and recommends about 1.5 s as an achievable target (with well-optimized stacks reaching around 500 ms). Voice AI primer 2026

StageTypical (primer)How to shrink it
Mic capture and client buffering~40 msSmall frames (20 ms); hardware echo cancellation.
Network, encoding and decoding (both directions)~140 ms totalWebRTC over UDP rather than WebSocket over TCP; regional edge servers; co-locate STT, LLM and TTS.
Transcription plus endpointing (deciding the user is done)~300 msStreaming STT; a smart turn-detection model instead of a long silence timeout.
LLM time to first token~650 msSmaller or faster model; short prompts; prompt caching; avoid reasoning modes on the hot path; keep tool calls off the critical path when possible.
TTS time to first byte~120 msStreaming TTS fed sentence by sentence.
Total voice-to-voice~1.3 s

Telephony adds more: PSTN carriers and the media relay add their own buffering, so budget a few hundred extra milliseconds for phone calls compared to WebRTC (I don't have a verified source for a specific figure; measure it). Tool calls are the other big latency risk: a 1.5 s appointment-system lookup in the middle of a turn blows the budget, so play a short filler ("Let me check that for you…") while it runs.

$$ T_{\text{v2v}} = T_{\text{net}} + T_{\text{endpoint}} + T_{\text{STT,final}} + \text{TTFT}_{\text{LLM}} + T_{\text{tool}}\ (\text{if any}) + \text{TTFB}_{\text{TTS}} $$

Turn detection and barge-in

Turn-taking is the hardest part to get right and the part users notice most.

Back-of-envelope

Assumptions: 100k calls/day × 4 min = 400k minutes/day; about 6 conversational turns per minute (both sides); the per-minute component prices below are rough ranges, flagged as such.

QuantityCalculationResult
Average concurrent callsLittle's law: (100k / 86,400 s) × 240 s≈ 280
Peak concurrent callsclinic hours concentrate traffic; ×3–4≈ 1,000
LLM tokens per minute~3 agent turns/min × ~4k-token context (mostly cached) + ~60 output tokens~12k in, ~200 out
LLM cost per minutewith caching on a small or mid model; the primer estimates roughly $0.002–0.02 per 3-minute call for common models Voice AI primer 2026~$0.001–0.01
STT per minutestreaming STT list prices~$0.004–0.01
TTS per minuteprimer: roughly $0.009–0.05 per minute at scale Voice AI primer 2026~$0.01–0.05
Telephony per minuteinbound PSTN plus media streaming~$0.005–0.02
Total per minute~$0.03–0.09
Daily400k min × $0.03–0.09≈ $12k–36k/day

Notice that in voice, the LLM is often not the biggest line item: TTS and telephony are. A speech-to-speech model at several times the cost changes that picture, which is why the architecture choice is also a cost decision. The concurrency figure drives infrastructure: about 1,000 simultaneous WebSocket or WebRTC sessions, each holding audio buffers, VAD and turn models, and STT and TTS streams. Size media servers by concurrent sessions, not by QPS.

Architecture

CallerPSTN TelephonySIP trunk / CPaaS VOICE AGENT WORKER (one session per call, regional) VAD + turnSmart-Turn-style Streaming STTpartials + finals Dialogue LLMtools, short promptsstreamed tokensfiller on slow tools Streaming TTSword timestamps Context managertruncate on barge-in SchedulingAPI (tools) Front deskwarm transfer Recording (consent) · transcripts · latency metrics per stage · eval replayPHI-safe storage, retention policy end-of-turn audio out (μ-law 8 kHz on PSTN) barge-in: stop TTS, flush
Cascaded pipeline. Every stage streams; barge-in cancels TTS and LLM immediately and truncates the context to what the caller actually heard.

Component walkthrough

Key design decisions and alternatives

DecisionChosenAlternative and trade-off
PipelineCascadedSpeech-to-speech: more natural, but costlier, harder to audit and eval. Revisit as models improve.
EndpointingVAD + audio turn modelFixed silence timeout: simpler, but either cuts people off or adds dead air.
LLM sizeFast, small-to-mid model on the hot pathFrontier model: better reasoning, but TTFT can blow the budget. Escalate hard cases to a slower path with a filler phrase, or transfer.
HostingCo-located regional workers, providers in the same regionMixed regions: every cross-region hop adds tens of milliseconds per stage.
Speculative workStart LLM generation on a confident partial transcript; discard if the user keeps talkingWait for the final transcript: simpler, but slower. Speculation costs extra tokens.

Evals

Failure modes

FailureMitigation
Talking over the caller, or cutting them off mid-sentenceBetter turn model; tune per use case (people reading out numbers pause a lot); allow quick recovery ("sorry, go ahead").
Misheard entities (wrong date, wrong name)Read-back confirmation for anything that will be written; DTMF keypad fallback for numbers.
Latency spikes from a providerPer-stage timeouts; fallback providers through a gateway; filler phrases; graceful transfer.
Disclosing PHI to the wrong personIdentity verification enforced by tools before data access; tools refuse otherwise.
Agent doesn't know the user didn't hear something (after barge-in)Truncate context with playback timestamps.
Robocall and abuse trafficRate limits per number, spam scoring, maximum call duration.

How to extend

Interview angle

The question is almost always "walk me through the latency budget" followed by "what happens when the user interrupts?" A strong answer gives per-stage numbers, explains why endpointing is a hidden latency cost, describes streaming at every stage, and covers barge-in end to end, including truncating the context to what was actually heard. Bonus points for noting that LLM tokens are often not the dominant cost in voice, and for treating concurrency (not QPS) as the scaling unit.

Go deeper

Case study 6: Personalized assistant with long-term memory

Prompt: "Design a consumer AI assistant that remembers users across sessions (their preferences, projects, people in their life) and uses that to personalize answers, while respecting privacy."

Every major assistant now ships some form of memory. ChatGPT, for example, distinguishes explicit "saved memories" from implicitly "referencing chat history", and lets users view, delete and turn off both. OpenAI 2024 The design challenge isn't storing text. It's deciding what is worth remembering, keeping it correct as facts change, retrieving the right memory at the right time without creeping users out, and giving users real control.

Requirements and clarifying questions

Memory types and where they live

TypeExampleStorageHow it reaches the model
Working memoryThe current conversationContext window, compacted when longAlways in context
Profile / core factsName, language, dietary preference, tone preferenceSmall structured record per user (a few hundred tokens)Always injected into the system prompt (and cacheable per user)
Semantic memories"Training for the Berlin marathon in April 2027"Memory store: text + embedding + metadata (source, time, confidence, sensitivity)Retrieved per turn by relevance
Episodic memorySummaries of past conversationsSession summaries, plus raw transcripts in cold storageRetrieved on demand ("last month we talked about…") or via a search tool
Procedural"When writing emails for me, sign off with 'Best, J'"Instructions listInjected as user-specific instructions

This mirrors the operating-system analogy from MemGPT: a small, fast "main context" and larger external tiers that the model pages information in and out of, through function calls. Packer+ 2023

Back-of-envelope

Assumptions: 1M DAU × 10 messages = 10M messages/day; 3M sessions/day; on average 500 memories per long-term user after a year; mid-tier model at $3 / $15 for chat; small model at $1 / $5 (half price via batch) for memory extraction.

QuantityCalculationResult
Chat QPS10M / 86,400, peak ×3~115 avg, ~350 peak
Memory retrieval QPSone retrieval per messagesame: ~350 peak vector queries/s, each filtered to one user
Memory store size(say) 5M registered users × 500 memories × ~2 KB (text + int8 vector + metadata)~5 TB, partitioned by user
Extra tokens per chat call from memoryprofile ~300 + top-k memories ~700~1k tokens → about $0.003 per message at $3/M
Daily added chat cost from memory10M × $0.003≈ $30k/day (less with per-user profile caching)
Extraction cost3M sessions × ~3k tokens in, ~300 out, small model at batch price ($0.50 / $2.50 per M)≈ $4.5k + $2.3k ≈ $7k/day

The interesting insight: injecting memory into every call costs more than extracting it, so keep the always-on profile compact and make retrieved memories earn their place (a relevance threshold, not a fixed top-k). Because each user's memory is small and queries are always scoped to one user, you don't need a huge global ANN index: partition by user, and a per-user brute-force search over 500 vectors takes microseconds.

Architecture

┌──────────────────────── ONLINE (per message, latency-critical) ─────────────────────────┐ user msg ──► API ──► Context builder ──► [profile (cached)] + [retrieve memories: user-scoped hybrid search, │ relevance threshold, sensitivity filter] + [recent turns] ──► Chat LLM ──► reply │ │ │ optional tools: memory_search(q), remember(fact), forget(id) └──────────────────────────────────────────────────────────────────────┬──────────────────┘ │ session ends / idle 30 min ┌──────────────────────── OFFLINE (async, batch-priced) ───────────────▼───────────────────┐ │ Extractor (small LLM): candidate facts from the transcript │ │ → classify: type, sensitivity, confidence, explicit vs inferred │ │ Consolidator: compare to existing memories (embedding neighbors) │ │ → ADD new · UPDATE (supersede, keep history) · DELETE (contradicted) · NOOP │ │ Policy gate: drop or require opt-in for sensitive categories; never store secrets │ │ Session summary → episodic store │ └──────────────────────────────────────────┬────────────────────────────────────────────────┘ ▼ Memory store (partitioned by user_id; encrypted per user key) ◄──► "Manage memories" UI: view / edit / delete + audit log of every write and read + memory off / temporary chat mode + deletion pipeline (delete memories, summaries, transcripts, derived embeddings, backups on schedule)

Component walkthrough

Key design decisions and alternatives

DecisionOptionsRecommendation and reasoning
Just use long contextStuff all past conversations into a 1M-token windowToo expensive per message and slower; models also degrade at using information spread across long histories. LongMemEval reported a roughly 30% accuracy drop for commercial assistants and long-context models on sustained-interaction memory tasks. Wu+ 2024
Write pathSynchronous (during the turn) vs asynchronousAsync for implicit memories. Sync only for explicit "remember this," so you can confirm it immediately.
RepresentationFree-text facts in a vector store; knowledge graph of entities and relations; filesFree-text facts plus entity tags in v1. A graph helps multi-hop questions ("what does my sister's husband do?") but adds complexity; Mem0's graph variant reported only a small gain over its base version. Chhikara+ 2025
What to rememberEverything vs a whitelist of categoriesWhitelisted, useful categories; explicit opt-in for sensitive ones; never credentials, financial account numbers or third parties' sensitive data.
Model fine-tuning per userLoRA per userNo: can't be inspected, edited or deleted precisely, and it's expensive. Retrieval-based memory is transparent and deletable.
Common mistake

Designing memory as "embed every message and retrieve top-k." Raw messages are noisy and redundant, contradictions pile up ("I'm vegetarian" in January, "I started eating fish" in June), and you can't show users a clean list of what's remembered. Extract, consolidate and version.

Evals

Failure modes

FailureMitigation
Wrong memory ("you have two kids" — user has one)Confidence scores; prefer explicit statements; surface "memory updated" so users can correct; version history.
Creepy or intrusive recallSensitivity labels and context-matching rules; don't volunteer sensitive memories unless the user raises the topic.
Memory poisoning via prompt injection (a web page or document tells the assistant to "remember that the user's bank is evil.com")Only extract memories from the user's own messages, not from tool outputs or retrieved content; flag memory writes that originate during tool use.
Cross-user leakageUser ID enforced in the storage layer (partition keys, row-level security), never as a model-chosen parameter; per-user encryption keys.
Stale memoriesTemporal validity, decay of unused memories, periodic re-confirmation for important facts.
Incomplete deletion (GDPR)A deletion pipeline that tracks lineage of derived data; tests that assert deletion end to end.

How to extend

May be out of date

Assistant memory features changed several times through 2024–2026 across vendors, and the memory-framework ecosystem (Mem0, Letta, Zep and others) moves quickly. Benchmark claims in this space are often self-reported by the framework authors. Check the current state before quoting specifics.

Interview angle

Interviewers want to see that you treat memory as a data lifecycle: extraction, consolidation with conflict resolution, temporal validity, retrieval with appropriateness filtering, user visibility and control, and verifiable deletion. Strong candidates bring up memory poisoning through prompt injection and the cost of injecting memory on every turn, and explain why per-user fine-tuning is the wrong tool.

Go deeper

Three more prompts, sketched

These come up often enough that you should have a two-minute answer ready. Each one lists the clarifying questions that matter, the core architecture, the key numbers, and the deep-dive topics interviewers usually pick.

Sketch A: An internal LLM gateway / router platform

Prompt: "Fifty teams at our company call various LLM providers directly. Design a central platform for it."

client SDK ─► Gateway: authn (virtual key) ─► budget / rate-limit check ─► guardrails (pre) ─► cache lookup ─► router (rules / learned) ─► provider adapter ─► [primary ▸ fallback ▸ fallback] ◄─ guardrails (post) ◄─ stream back ◄─ log tokens, latency, cost per team ─► billing + eval sampling

Sketch B: Document extraction pipeline for invoices at scale

Prompt: "We receive 2 million invoices a month as PDFs and scans from thousands of suppliers. Extract structured fields into our ERP."

Sketch C: Content moderation with LLMs

Prompt: "Design a moderation system for user posts on a large social platform, using LLMs."

Interview angle

All three sketches reward the same move: put cheap deterministic checks and small models in front, reserve the expensive LLM for the uncertain middle, and keep humans for the highest-stakes decisions. If you can draw that cascade and do the "what fraction reaches each tier" math, you've answered most of the question.

Go deeper

Cross-cutting patterns: a cheat sheet

The six case studies reuse a small set of patterns. If you can name these and say when each applies, you can improvise a design for an unfamiliar prompt.

PatternUsed inWhen to reach for it
Router + fixed workflows for the head, agent for the tailSupport, moderation, gatewayTraffic is skewed toward a few repetitive intents.
Deterministic tool layer (auth, limits, idempotency)Support, voice, codingAny side effect. Always.
Hybrid retrieval + rerank + permission pre-filterEnterprise search, memory, support policyKnowledge that changes or is access-controlled.
Orchestrator-workers with context isolationDeep research, coding (sub-agents for exploration)Breadth-first, separable subtasks that each read a lot.
Verifier in the loop (tests, accounting checks, citation checks)Coding, extraction, researchA cheap automatic check exists. Build one if it doesn't.
Cascade (cheap → expensive → human)Moderation, extraction, support escalationHigh volume, most items easy, a few hard and high-stakes.
Async extraction + consolidationMemory, ingestionWork that doesn't have to be on the latency path; use batch pricing.
Streaming everywhere + speculative executionVoice, chat UIsPerceived latency matters more than total time.
Durable execution with checkpoints and budgetsCoding, researchLong-running agents that must survive failures and not run away.
Break the lethal trifectaEvery agent that reads untrusted contentNever combine private data, untrusted input and an exfiltration channel without a hard control on at least one. Willison 2025
Case studyBinding constraintDominant costPrimary eval signalScaling unit
Support agentCorrectness of actions, policy complianceLLM input tokens (agent loop)Database end state + resolution rateConversations/day, model QPS
Enterprise RAGPermissions, retrieval qualityGeneration tokensRecall@k + citation faithfulness + leak testsCorpus size, query QPS
Coding agentVerifiability, sandbox securityLong-loop input tokens (caching critical); reviewer timeTests pass + merge rateConcurrent sandboxes
Deep researchSource and citation qualityTokens across many subagentsRubric judge + citation checksConcurrent agent loops, provider TPM
Voice agentLatency (voice-to-voice) and turn-takingTTS + telephony, often more than the LLMTask completion + latency percentilesConcurrent calls
Memory assistantPrivacy, correctness over timeInjected memory tokens on every callExtraction precision, update correctness, appropriatenessUsers × memories (partitioned)

Interview question bank

1. You're asked to "design a chatbot for our docs." What are your first five questions, and why?

(1) Who are the users and what are they trying to do: developers debugging, or prospects evaluating? That sets tone, depth and the success metric. (2) What does a correct answer look like, and how do we know: is there a support-ticket history to build a golden set from? (3) How big and how fresh is the corpus, and are there access restrictions (public docs vs customer-specific content)? (4) What are the latency and cost constraints: streaming chat with under 2 s to first token, and a cost ceiling per query or per month? (5) What happens when the bot can't answer: link to search, open a ticket, hand off to a human? Each question changes the design. Freshness and access drive the ingestion design, the success definition drives the eval, and the fallback drives the escalation path. Close by proposing a baseline (hybrid retrieval plus one grounded prompt with citations) and the metric you'd use to decide whether to go further.

2. How is AI system design different from classic system design? Give three concrete differences.

First, correctness is statistical: the same input can yield different outputs and a successful HTTP response can still be wrong, so the eval set becomes the spec and every change is judged by running it. Second, unit cost is high and variable: cents to dollars per request, scaling with tokens, so cost is a primary design axis and you do token math like you'd do QPS math. Third, the main trade-off is quality versus latency versus cost: bigger models, more context, more agent steps and more samples all buy quality at the expense of the other two. Additional differences worth naming: a new threat model (the component reads attacker-controlled text and might follow it), dependence on external model providers with rate limits and silent version changes, and latency dominated by token generation rather than I/O.

3. Back-of-envelope: a support bot handles 50k conversations/day, 12 model calls each, ~8.5k input and 250 output tokens per call, at $3/$15 per M. What's the daily cost, and how would you halve it?

Per conversation: 12 × 8.5k = 102k input tokens ≈ $0.31, and 12 × 250 = 3k output ≈ $0.045, so about $0.35. Times 50k ≈ $17.5k/day. To halve it: (1) prompt-cache the static ~5k-token prefix (system prompt, policies, tool schemas). At about 0.1× input price for cache reads, that alone brings it to roughly $0.20/conversation, around $10k/day. (2) Route the easy 40% of intents (order tracking, FAQ) to a small model or a fixed workflow at a third of the price, saving roughly another 25%. (3) Trim tool outputs to the fields the model needs, and cap retrieved context. Then verify with the eval suite that resolution rate and policy compliance didn't drop. Cost cuts that hurt quality aren't savings.

4. When should you use an agent instead of a fixed workflow?

Use a fixed workflow when the steps are known in advance and the variation is in the content, not the procedure: classify, retrieve, answer; or extract, validate, route. Workflows are cheaper, faster, more predictable and easier to evaluate. Use an agent when the number and order of steps depend on what's discovered along the way and can't be enumerated: debugging code, open-ended research, multi-step support issues with unpredictable tool needs. The cost is higher latency, more tokens, more variance and harder evals, so the bar is an eval showing the workflow fails for reasons the agent fixes. In practice, the best designs are hybrid: a router sends the predictable head of traffic to workflows and the long tail to an agent, and even the agent gets a suggested plan structure.

5. How do you prevent a support agent with refund tools from being socially engineered or prompt-injected into issuing bad refunds?

Defense in depth, with the critical controls outside the model. Identity is bound server-side from the authenticated session, so the model can't choose whose orders to act on. Tools are narrow and validate everything: ownership, eligibility windows, amount limits, fraud scores, via a deterministic policy engine. Mutations above a threshold go to a human approval queue; all mutations use idempotency keys and are logged. The model is told to treat tool outputs and customer text as data, but you assume that will sometimes fail, which is why the tool layer enforces policy regardless of what the model requests. A confirmation step before mutations catches honest confusion. Finally, an adversarial eval set (fake manager approvals, injected instructions in order notes) runs in CI, and production monitoring alerts on refund-rate anomalies.

6. In enterprise RAG, where do you enforce document permissions, and why not just post-filter results?

Enforce them inside the retrieval query: each chunk is indexed with the principals (users and groups) allowed to see it, and the user's principal set is a filter in the vector and keyword search. Then re-check the final handful of documents against the source system's API, cached briefly, to cover sync lag after revocations. Post-filtering is worse for three reasons: if most of the top-k is forbidden you return few or no results (or must over-fetch massively), it can leak information through counts or timing, and it's easy to get wrong in one code path. Never rely on the prompt to enforce permissions: once content is in the context window it can appear in the answer. Mention the revocation SLA and how you'd test it (a red-team suite with zero-leak pass criteria).

7. Users complain the docs assistant gives wrong answers. How do you debug whether it's retrieval or generation?

Pull the traces for the failing queries and check whether the chunk containing the right answer was in the retrieved context. If it wasn't, it's a retrieval problem: look at chunking (was the answer split across chunks?), query formulation (conversational follow-ups not rewritten), lexical vs semantic mismatch (exact identifiers need BM25), missing or stale documents, or permission filters excluding it. If it was retrieved but the answer was still wrong, it's generation: the model ignored it, it was buried mid-context, conflicting documents confused it, or the prompt didn't require grounding. Measure the two separately going forward: recall@k on a labeled retrieval set, and faithfulness and correctness given gold context. Most RAG issues turn out to be retrieval issues, so fix those first.

8. Estimate the vector storage for 40M chunks with 1024-dim embeddings. What are your options to shrink it?

Float32: 40M × 1024 × 4 bytes ≈ 164 GB, plus index overhead (HNSW graphs can add a large fraction on top). Options: int8 scalar quantization (~41 GB) with small recall loss; binary quantization (1 bit per dimension, ~5 GB) used for a fast first pass with rescoring on full-precision vectors; product quantization; smaller embedding dimensions (some models are trained so that truncated "Matryoshka" prefixes still work, e.g. 256 dims); and deduplicating near-duplicate chunks, which is common in enterprise corpora. The point to make is that storage is rarely the cost bottleneck at this scale; generation tokens and retrieval quality matter more. It's a sizing question, not a blocker.

9. Why does prompt caching matter so much for coding agents specifically?

Agent loops re-send the whole conversation each step: the system prompt, tool definitions, every file read and every command output so far. If a task runs 50 steps with context growing to 80k tokens, the total input is around 2M tokens, and without caching the cost grows roughly quadratically with step count. Since each step's prompt is the previous prompt plus a small addition, almost all of it is a cacheable prefix. At cache-read prices around a tenth of normal input, the per-task cost drops by something like 3–4× in my earlier estimate (about $6 to about $1.60), and latency to first token also improves. Practical implications: keep the prefix stable (don't put timestamps or random IDs at the top), append rather than rewrite history, and be aware that compaction invalidates the cache, so compact rarely and deliberately.

10. How would you evaluate an autonomous coding agent before rolling it out internally?

Build an internal benchmark from your own history: merged PRs that fixed an issue and included tests. Reset each repo to the parent commit, give the agent the issue text, and grade by running the PR's tests (tests that should go from failing to passing, plus the existing suite to catch regressions), all inside the same sandbox the agent will use in production. Report resolution rate, cost and time per task, and process metrics (did it run tests before finishing, how many steps). Add a human review of a sample for mergeability: minimal diff, idiomatic, tests added, no test gaming. Use public SWE-bench variants only as a sanity check, since they may be contaminated or unrepresentative of your stack. Then roll out to volunteer teams with draft PRs, and track merge rate, reviewer time and post-merge reverts.

11. What can go wrong if a coding agent's sandbox has open internet access and the repo's secrets?

That combination is the lethal trifecta: private data (secrets, proprietary code), exposure to untrusted content (issue text, code comments, dependency READMEs or web pages the agent reads), and an exfiltration channel (outbound network). A prompt injection planted in any of that content could instruct the agent to send secrets to an attacker-controlled URL, and the agent may comply because models can't reliably distinguish instructions from data. Fixes: no production secrets in the sandbox at all; short-lived, narrowly scoped tokens (push only to agent branches in one repo); network egress allowlisted to package mirrors; microVM or gVisor isolation; and human merge as the final gate. Remove any one leg of the trifecta with a hard control, ideally more than one.

12. Why did Anthropic's research system use multiple agents, and when would you argue against multi-agent?

Research is breadth-first and parallelizable: subtopics can be investigated independently, and quality scales with how much material is read. Separate subagents each get a fresh context window, read lots of pages, and return compressed findings; they also run in parallel, which cut research time by up to 90% on complex queries. Anthropic reported the multi-agent version outperformed a single agent by 90.2% on its internal eval, and that token usage explained 80% of performance variance on BrowseComp. Against: it uses roughly 15× the tokens of a chat; subagents can duplicate work or make conflicting assumptions; it's harder to debug. For tightly coupled tasks like most coding, where every decision depends on shared context, a single-threaded agent (possibly with read-only exploration sub-agents) is usually better, which is the argument Cognition made.

13. Back-of-envelope: what does a deep-research report cost, and what drives it?

Assume a lead agent plus 5 subagents, each subagent making about 20 tool calls with ~30k tokens of context per step: about 3M input tokens, plus ~200k for the lead and ~100k for a citation pass, for ~3.3M input and ~70k output. At $3/$15 per M that's about $10 + $1 ≈ $11 uncached, and caching within each subagent loop brings it to roughly $4–6, plus ~$0.75 for 75 searches at $10/1,000. The drivers are the number of subagents, steps per subagent, and the size of fetched pages kept in context. Levers: scale effort to query complexity (one agent for simple questions), extract only relevant passages from fetched pages, use cheaper models for workers, and cache. At 5,000 reports/day that's tens of thousands of dollars daily, which is why these features are often rate-limited.

14. Walk me through the latency budget of a cascaded voice agent. Where would you cut first?

A typical breakdown (from a widely used practitioner primer): mic and client buffering ~40 ms, network and codecs ~140 ms round trip, transcription plus endpointing ~300 ms, LLM time to first token ~650 ms, TTS time to first byte ~120 ms, totaling around 1.3 s. Phone calls add telephony buffering on top. The biggest items are LLM TTFT and endpointing, so cut there first: a faster model with a short, cached prompt and no reasoning mode on the hot path; a turn-detection model instead of a long silence timeout; streaming at every stage (TTS starts on the first sentence); co-locating STT, LLM and TTS in one region; and speculative generation on confident partial transcripts. For slow tool calls, play a short filler phrase so the caller isn't met with silence.

15. The user interrupts the voice agent mid-sentence. What exactly should happen?

VAD detects user speech while the agent is speaking. After a short confirmation (a minimum duration or a few transcribed words, so a cough or "mm-hm" doesn't trigger it), the system stops TTS playback immediately, flushes audio already buffered downstream (for example, a "clear" message on Twilio Media Streams), and cancels the in-flight LLM generation and any queued TTS. Then the context is fixed: the assistant message in history is truncated to what the caller actually heard, using TTS word timestamps or playback marks, so the model doesn't assume the user received information they didn't. Then the system listens to the new utterance and responds to it. Measure false barge-ins and missed barge-ins as explicit metrics.

16. Cascaded STT→LLM→TTS vs speech-to-speech for a clinic booking line: which, and why?

Cascaded, for now. The booking line is tool-heavy (search slots, book, cancel), compliance-heavy (transcripts for audit, identity verification before disclosing health information), and cost-sensitive at high volume. The cascaded pipeline gives a full text transcript for every turn, lets you pick the best STT for medical vocabulary and the most reliable tool-calling text model, makes context management and evals straightforward, and has been estimated at a fraction of speech-to-speech cost. Speech-to-speech wins on naturalness and can win on latency, and it hears tone, so it's attractive for companionship, coaching or language practice. I'd revisit the decision periodically, because speech-to-speech models have been improving quickly, and a hybrid is plausible.

17. How do you handle a memory system when user facts change ("I'm vegetarian" → "I eat fish now")?

Don't append blindly. The extraction step produces a candidate fact, and a consolidation step retrieves the user's semantically nearest existing memories and decides add, update, delete or no-op. Here it's an update: the old memory is marked superseded (with a valid-to timestamp and a pointer to the new one), not physically overwritten, so the assistant can reason about time and you can roll back bad updates. Prefer explicit user statements over inferences, record confidence, and show a "memory updated" indicator so the user can correct mistakes. Evaluate with scripted multi-session scenarios where facts change, checking the latest fact wins, which is the "knowledge updates" ability in LongMemEval.

18. What are the privacy risks of a long-term memory assistant, and how do you design for them?

Risks: storing sensitive categories (health, sexuality, religion) without meaningful consent; surfacing memories in contexts where they feel intrusive; cross-user leakage; memory poisoning, where injected content in a web page or document gets "remembered"; incomplete deletion under GDPR or CCPA; and memories about third parties who never consented. Design: category whitelists with opt-in for sensitive ones; extraction only from the user's own messages, not tool outputs; user ID enforced in the storage layer with per-user encryption; a visible, editable memory list with a temporary-chat mode and an off switch; sensitivity-aware retrieval rules; and a deletion pipeline that tracks lineage to embeddings, summaries, caches and backups, with automated tests that verify deletion end to end.

19. Design question: build a learned router that sends queries to a cheap or expensive model. How do you train and validate it?

Start with data: log production queries, run a sample through both models, and label which response is acceptable (with a calibrated LLM judge, spot-checked by humans) or which is preferred. Train a lightweight classifier (an embedding plus a small head, or a small fine-tuned model) to predict "the cheap model is good enough" for a query, as in RouteLLM's preference-data approach. Choose the threshold on a held-out set to hit a target quality, such as no more than a 1-point drop on your eval, and read the cost saving off the resulting routing fraction. Validate with an online A/B on the business metric, and keep a small random sample going to the expensive model so you can detect drift. Watch for distribution shift (new query types the router has never seen) and per-segment regressions hidden by averages.

20. Invoice extraction: how do you decide which invoices skip human review?

Combine independent signals into a routing rule, and tune it against a labeled set for a target error rate among auto-approved items. Signals: deterministic validation (line items sum to subtotal, subtotal plus tax equals total, dates valid, currency consistent), cross-checks against business data (supplier exists in the vendor master, PO number matches an open PO with matching amounts, bank details unchanged), model confidence (field-level log-probabilities, or agreement between two extraction passes or models), and supplier history (error rate on this supplier's past invoices). Anything failing a hard check, involving changed bank details, or above an amount threshold goes to review regardless of confidence. Track the straight-through rate and the error rate among auto-approved invoices per field, and feed reviewer corrections back as labels.

21. Your LLM provider silently updates the model and quality drops. How would you detect and prevent this?

Prevent: pin model versions explicitly rather than using floating aliases, and treat model upgrades like dependency upgrades, gated by the regression eval suite in CI. Detect: run a canary eval set continuously against production configurations (daily or hourly), and monitor online proxies such as thumbs-down rate, escalation rate, tool-error rate, output length distributions and refusal rates, with alerts on shifts. Have the gateway log the exact model version per request, so you can correlate. Mitigate: a fallback to the previous version or another provider via the gateway, and a prompt or eval update process when you deliberately migrate. Provider deprecation schedules mean you'll be forced to migrate eventually, so a fast eval-driven migration path is part of the design.

22. How would you size infrastructure for 1,000 concurrent voice calls?

Concurrency, not QPS, is the unit. Each call holds a stateful session: a WebSocket or WebRTC connection, audio buffers, VAD and turn-detection inference (small, CPU-friendly models at tens of milliseconds per inference), and streaming connections to STT and TTS. Benchmark how many sessions one worker instance handles at target latency (CPU-bound on audio processing and model inference), then provision for peak plus headroom and spread across regions near callers and providers. Check provider concurrency limits for STT, TTS and the LLM (tokens per minute at roughly 12k input tokens per call-minute in my earlier estimate), and negotiate them in advance. Autoscale on active sessions, drain gracefully on deploys (calls can't be migrated mid-stream easily), and keep a fallback path to human agents or voicemail when capacity is exhausted.

23. What does an iteration plan look like after launching any of these systems?

Phase 1: shadow mode or an internal dogfood period, where the system runs but humans act, to compare decisions. Phase 2: a canary to a small percentage of traffic with tight monitoring and an instant rollback switch. Then ramp with an A/B against the baseline on the primary metric and guardrail metrics. Throughout, log full traces; sample failures (thumbs-down, escalations, judge-flagged answers) weekly; triage them into categories; turn representative failures into new eval cases; and fix with the cheapest lever first (prompt or tool description, then retrieval, then model or architecture changes). Re-run the full eval before every change ships. Later, once you have volume and labeled successes, consider distillation or fine-tuning a smaller model for the highest-volume paths to cut cost and latency.