AI System Design Case Studies
The "design an AI system" round is where everything else in this guide gets tested together: retrieval, agents, evals, serving, cost and safety. Interviewers aren't looking for a perfect diagram. They want to see that you can turn a vague product ask into a measurable task, start from the simplest thing that could work, do the token and dollar math out loud, and design for a component that is non-deterministic and sometimes wrong. This page gives you a reusable framework, then works through six full case studies and three short ones in the shape a strong candidate would present them.
TL;DR: the 8–12 things to be able to say out loud
- Clarify the task as an eval before drawing boxes. "What does a correct output look like, who judges it, and what's the cost of a wrong one?" Everything downstream depends on this.
- Start with a one-prompt baseline, then add retrieval, tools, multiple steps or multiple agents only when a measured failure justifies it.
- The quality / latency / cost triangle is the core trade-off. Every lever (bigger model, more context, more steps, more samples) buys quality with latency and money.
- Cost per request = tokens × price, and agent loops multiply tokens because the context is re-sent every step. Prompt caching (cache reads at roughly a tenth of the input price) and model routing are the first two cost levers.
- Agents re-read their whole context every turn, so input tokens dominate cost; output tokens dominate latency.
- Permissions and side effects live in deterministic code, not in the prompt. Tools enforce auth, limits and idempotency; the model only proposes.
- Evals are the spec: offline golden sets plus LLM-as-judge with calibrated rubrics, outcome-based checks (database state, tests passing), and online metrics (resolution, escalation, thumbs, cost).
- Prompt injection is the default threat model for any system that reads untrusted content and can act. Break the "lethal trifecta": private data + untrusted input + an exfiltration channel.
- Human-in-the-loop is a design component: confidence thresholds, approval gates for irreversible actions, and graceful hand-off with full context.
- Voice is a latency problem: about 1–1.5 s voice-to-voice is a realistic target, and turn detection matters as much as model speed.
- Memory is a data-governance problem as much as a retrieval problem: what to store, how to update or forget, and how the user sees and edits it.
- Close with an iteration plan: ship to a slice, log traces, mine failures into the eval set, and tune the cheapest lever first.
A reusable framework for AI system design interviews
Classic system design interviews follow a familiar arc: requirements, API, data model, high-level design, deep dives, scaling. AI system design keeps that arc but adds steps for the parts that are new: defining quality, choosing and composing models, evaluating a stochastic component, and paying per token. Here is a ten-step framework you can run through in roughly 40 minutes. You won't spend equal time on each step. The interviewer will usually pull you into two or three deep dives, so keep the rest brief and signpost what you're skipping.
| # | Step | What you actually say or produce | Time |
|---|---|---|---|
| 1 | Clarify the task and success metrics | Who the users are, the inputs and outputs, what "correct" means, and the business metric (resolution rate, time saved, conversion). Turn that into an eval definition. | 4–6 min |
| 2 | Constraints | Latency (p50/p95, streaming or not), cost ceiling per request or per month, scale (DAU, QPS, corpus size), privacy and residency, the accuracy bar and the cost of errors, regulatory needs. | 2–3 min |
| 3 | Baseline | The simplest thing that could work: one prompt to a capable model with the obvious context pasted in. Say what it would get right and where it would fail. | 2 min |
| 4 | Data and context sources | What the model needs to know (docs, DB rows, user history, tools), freshness, access control, volume. Retrieval vs tools vs fine-tuning. | 3–5 min |
| 5 | Architecture | Diagram: entry point, orchestrator, retrieval, tools, model calls, state store, guardrails, human hand-off, logging. Workflow (fixed steps) vs agent (model picks the steps). | 8–10 min |
| 6 | Model choices | Which tier per step (small/fast for routing and classification, frontier for hard reasoning), hosted API vs open weights, fine-tune or not, structured outputs. | 2–3 min |
| 7 | Evals | Offline golden set, LLM judge plus rubric, outcome checks, regression gates in CI, online metrics and A/B, human review sampling. | 4–6 min |
| 8 | Failure modes and safety | Hallucination, wrong tool call, prompt injection, data leakage, runaway loops, provider outage. A mitigation for each one. | 3–5 min |
| 9 | Scaling and cost math | QPS, tokens per request, dollars per day, GPU or rate-limit needs, caching, batching, routing. | 3–5 min |
| 10 | Iteration plan | Rollout (shadow, canary, % ramp), the feedback loop from traces to the eval set, what you'd build in v2. | 2 min |
Step 1 in more depth: turn the ask into an eval
The most common weak opening is jumping to "we'll use RAG with a vector DB." The strong opening is a set of questions that pin down the task as something you could grade:
- Unit of work: a single answer, a multi-turn conversation, a completed action (refund issued), or an artifact (a PR, a report)?
- Ground truth: is there a checkable answer (database state, tests pass, extracted field equals the label), or is quality subjective (helpfulness, tone), which needs rubrics and judges?
- Error asymmetry: what does a false positive cost compared to a false negative? Wrongly issuing a $500 refund is different from wrongly escalating to a human.
- The business metric and the guardrail metrics: for example, raise automated resolution rate while holding CSAT and refund-loss rate flat.
Then say it as a sentence: "We'll call v1 successful if, on a golden set of 500 real tickets, it fully resolves at least 60% with no policy violations, escalates the rest with a correct summary, keeps p95 first-token latency under 2 s, and costs under $0.30 per conversation." The numbers are placeholders you agree with the interviewer, but having them changes the whole conversation.
Step 3 in more depth: why the baseline matters
Anthropic's guidance on agents is to find the simplest solution and add complexity only when it demonstrably helps. Their taxonomy separates workflows (LLMs and tools orchestrated through predefined code paths: prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer) from agents (the model dynamically directs its own process and tool use). Anthropic 2024 Interviewers reward candidates who climb that ladder explicitly:
A related data point from coding: the "Agentless" paper showed that a fixed three-phase pipeline (localize, repair, validate) beat the open-source agents of the time on SWE-bench Lite, at about $0.70 per issue. Xia+ 2024 Agents have improved a lot since then, but the lesson holds: a well-structured workflow is often a strong baseline, and you should be able to argue for it.
How AI system design differs from classic system design
| Dimension | Classic system design | AI system design |
|---|---|---|
| Correctness | Deterministic. A request either succeeds or errors. | Stochastic and graded. The same input can give different outputs, and "200 OK" can still be wrong. Quality is a distribution you measure. |
| Spec | API contract and requirements doc. | The eval set is the spec. You can't say a change is an improvement without running it. |
| Unit cost | Fractions of a cent per request; compute is amortized. | Cents to dollars per request, scaling linearly with tokens. Cost is a first-class design constraint, often the binding one. |
| Latency | Milliseconds, dominated by I/O and DB. | Seconds, dominated by token generation (roughly proportional to output length) and sequential model calls. |
| Main trade-off | Consistency / availability / latency. | Quality / latency / cost, plus safety. Bigger models, more context and more steps buy quality at the price of the other two. |
| Failure modes | Crashes, timeouts, data races. | Confident wrong answers, wrong tool arguments, prompt injection, runaway loops, silent quality regressions when a provider updates a model. |
| Testing | Unit and integration tests with exact asserts. | Statistical evals, LLM judges, pass@k and pass^k, human review sampling, online A/B. |
| Dependencies | Your own services. | Often an external model API with rate limits, outages, deprecations and price changes. A gateway and fallbacks are needed. |
| Security | Authn/z, input validation. | All of that, plus: the model reads attacker-controlled text and may follow it. Treat model output as untrusted input to your tools. |
What interviewers actually grade
Signals of a strong candidate
- Asks clarifying questions that change the design (error cost, scale, latency mode).
- States success metrics and an eval plan early.
- Proposes a simple baseline and justifies each added component.
- Does token and dollar math with stated assumptions, out loud.
- Puts auth, limits and irreversible actions in code, not prompts.
- Names specific failure modes with specific mitigations.
- Knows rough magnitudes: TTFT, tokens/s, price tiers, embedding sizes.
- Has an iteration and rollout plan; treats production as the data source.
Signals of a weak one
- Starts with "multi-agent system with a vector DB" before knowing the task.
- "We'll fine-tune" as a reflex, with no data plan.
- No numbers: no QPS, tokens or dollars.
- "The LLM will check permissions."
- Evals amount to "we'll test it."
- Ignores latency for user-facing paths, or cost for agents.
- No human fallback; no plan for when the model is wrong.
- Buzzword lists without trade-offs.
Back-of-envelope toolkit
You'll reuse the same handful of formulas in every case study. Memorize the shape, not the exact prices, which change every few months.
$$ \text{cost/request} = \sum_{\text{calls}} \Big( T_{\text{in,uncached}}\,p_{\text{in}} + T_{\text{in,cached}}\,p_{\text{cache}} + T_{\text{out}}\,p_{\text{out}} \Big) $$ $$ \text{QPS}_{\text{avg}} = \frac{\text{requests/day}}{86{,}400}, \qquad \text{QPS}_{\text{peak}} \approx (2\text{–}5)\times \text{QPS}_{\text{avg}} $$ $$ \text{latency} \approx \text{TTFT} + \frac{T_{\text{out}}}{\text{tokens/s}} \quad\text{(per sequential call)}, \qquad \text{concurrency} = \text{arrival rate} \times \text{duration} $$The last one is Little's law, and it's what sizes long-running agents and voice calls: 100 calls/s arriving that each last 4 minutes means about 24,000 concurrent calls.
| Rule of thumb | Value to use (state it as an assumption) |
|---|---|
| Tokens per English word | About 1.3 (roughly 4 characters per token). Newer tokenizers vary; Anthropic notes its Claude 4.7+ tokenizer produces about 30% more tokens for the same text. Anthropic pricing |
| Mid-tier hosted model list price (late 2026) | $2–3 per M input tokens, $10–15 per M output. On this page I use $3 / $15 as a round, slightly conservative number. |
| Small / fast model | About $1 / $5 per M (e.g. Haiku-class) Anthropic pricing |
| Frontier model | About $4–5 / $20–25 per M, with premium tiers above that |
| Prompt-cache read | About 0.1× the base input price on Anthropic's API (5-minute cache write costs 1.25×). Anthropic pricing Other providers have similar discounts with different mechanics. |
| Batch API discount | 50% for asynchronous jobs on Anthropic's API Anthropic pricing |
| Time to first token (hosted, short prompt) | 200–800 ms; grows with prompt length and reasoning |
| Output speed | 50–150 tokens/s for typical hosted models; specialized inference hardware can be much faster |
| Embedding vector | 1024 dims × 4 bytes = 4 KB float32; about 1 KB int8; 128 bytes binary |
Model prices dropped steeply through 2024–2026 and model names turn over every few months. The pricing figures above come from Anthropic's pricing page as of October 2026. In an interview, say "assuming roughly $3/$15 per million tokens" and move on. The interviewer cares about the structure of the calculation, not the third significant figure.
A common probe is "how would you cut the cost by 10×?" A strong answer ranks levers by effort and risk: (1) prompt caching for the static prefix, which is nearly free; (2) trimming context with better retrieval and shorter tool outputs; (3) routing easy traffic to a small model, measured by evals; (4) batching non-interactive work at 50% off; (5) distilling or fine-tuning a small model on logged traffic once you have volume. Then say what you'd measure to make sure quality held.
- Anthropic: Building effective agents: the workflow vs agent taxonomy and the "start simple" argument.
- Anthropic: Effective context engineering for AI agents: compaction, note-taking, sub-agents, just-in-time retrieval.
- OWASP Top 10 for LLM Applications (2025): a ready-made failure-mode checklist (prompt injection, excessive agency, unbounded consumption).
Case study 1: Customer-support agent with tool access and human escalation
Prompt: "Design an AI agent for an e-commerce company that handles customer support chat. It should be able to look up orders, process refunds and cancellations, and escalate to human agents when needed."
This is the most common AI design question because it touches everything: conversation, retrieval over policy docs, tool calling with side effects, authorization, escalation, and a clear business metric. Anthropic uses customer support as the canonical example of where agents fit, because it pairs conversation with actions and success is measurable by resolution. Anthropic 2024
Requirements and clarifying questions
| Question | Why it matters | Assumed answer |
|---|---|---|
| Channels: web chat only, or email and voice too? | Latency budget and turn structure differ a lot. | Web and in-app chat first; email later. |
| Is the user authenticated? | Determines whether the agent can act at all, and how identity is bound to tools. | Yes, logged-in users. Guests get FAQ only. |
| Which actions, and with what limits? | Side effects need policy limits and approval gates. | Order lookup, tracking, cancel before shipment, refunds up to $200 automatically; above that needs human approval. |
| Volume and peaks? | Sizing, rate limits, cost. | 50k conversations/day, 3× peak on sales days. |
| Current baseline? | Defines success. | Human agents; average cost per contact $4–6; CSAT 4.1/5. |
| Languages and regions? | Model choice, policy variants, data residency. | English and Spanish; US and EU (GDPR). |
| What does the hand-off look like? | Integration with the ticketing tool, agent desktop. | Escalate into the existing helpdesk queue with summary and transcript. |
Success metrics. Primary: automated resolution rate (conversation ends with the issue solved and no human touch, and no re-contact within 7 days). Guardrails: CSAT not below baseline, policy-violation rate (refunds outside policy) below 0.5%, wrong-action rate near zero, escalation quality (the human didn't have to re-ask). Operational: p95 time-to-first-token under 2 s, cost per conversation under $0.40.
Back-of-envelope
Assumptions (state them explicitly): 50k conversations/day; 6 user turns per conversation; an average of 2 model calls per user turn (one to decide a tool call, one to answer after the tool returns), so 12 calls per conversation; a mid-tier model at $3 / $15 per M tokens.
| Quantity | Calculation | Result |
|---|---|---|
| Model calls/day | 50k × 12 | 600k |
| Average QPS (model calls) | 600k / 86,400 | about 7 |
| Peak QPS | × 3 on sale days | about 20 |
| Input tokens per call | system prompt + policies 3k, tool schemas 2k, conversation so far ~2k avg, retrieved policy snippets 1.5k | ~8.5k |
| Output tokens per call | tool-call JSON or a short reply | ~250 |
| Per conversation, uncached | 12 × 8.5k = 102k in → $0.31; 12 × 250 = 3k out → $0.045 | ≈ $0.35 |
| Per conversation, with prompt caching | the 5k static prefix is cached on 11 of 12 calls: 55k tokens at $0.30/M ≈ $0.017; the rest (~47k) at $3/M ≈ $0.14; output $0.045 | ≈ $0.20 |
| Daily model spend | 50k × $0.20–0.35 | ≈ $10k–17.5k/day |
| Routing the easy ~40% (FAQ, tracking) to a small model at ~1/3 the price | another ~25% saving |
Compare to the human baseline: if 60% of 50k conversations are automated at roughly $0.25 each, that replaces about 30k human contacts a day. At $4–6 each, that's well over $100k/day of human cost against roughly $12k of model spend. The economics are rarely the blocker here; trust, correctness and policy compliance are. Say that out loud: it tells the interviewer where you'll spend design effort.
Architecture
Component walkthrough
- Gateway. Terminates the authenticated session and attaches a server-side customer ID. The model never sees or chooses whose data to access; the customer ID is bound to the tool credentials, not passed as a free-form argument.
- Input guard. A cheap classifier (small model or a dedicated safety model) flags abuse, self-harm, legal threats and obvious injection attempts, and redacts card numbers before anything is logged.
- Orchestrator. A small, fast model first classifies intent (tracking, return, refund, cancel, account, complaint, other). Simple intents run a fixed workflow: "where is my order" is one tool call and a templated answer. Open-ended issues go to the agent loop, where a mid-tier model with tools and policy context iterates. The orchestrator also holds a per-conversation state machine (identified order, eligibility checked, action taken) and a budget (max steps, max tokens, max tool calls).
- Policy RAG. Return and refund policies, shipping FAQs, and product-specific rules, chunked by section and versioned. Retrieval filters on region and product category so the agent never mixes EU and US return windows. Policies change, so each answer logs the policy version it used.
- Tool service. The most important box. Tools are narrow and intention-revealing (
get_order(order_id),check_refund_eligibility(order_id, items),issue_refund(order_id, items, reason, idempotency_key)) rather than a generic "call the Orders API." Every call is checked against a deterministic policy engine (eligibility windows, amount limits, the order belongs to this customer). Mutations take idempotency keys so a retried step can't double-refund. Anthropic's tool-design guidance makes the same points: design tools for the agent's affordances rather than wrapping raw APIs, namespace them, and return compact, high-signal results. Anthropic 2025 - Approval queue. Refunds above the auto-limit, or flagged by the fraud score, become a pending action that a human approves with one click. The agent tells the customer truthfully that it's been submitted for review.
- Escalation. Triggered by explicit request ("talk to a human"), low confidence, repeated failure (same intent three turns running), negative sentiment, or sensitive categories (legal, safety, chargebacks). The hand-off writes a structured summary (issue, order, what was tried, customer sentiment) so the human doesn't restart the conversation.
- Traces. Every prompt, retrieved chunk, tool call, tool result and decision is logged with a conversation ID. This is the raw material for evals and debugging.
Key design decisions and alternatives
| Decision | Chosen | Alternative | Why |
|---|---|---|---|
| Control flow | Hybrid: router + fixed workflows for the top intents, agent loop for the long tail | Pure agent for everything | Fixed flows are cheaper, faster and easier to eval for the 60–70% of traffic that's repetitive. The agent handles the messy remainder. |
| Where policy lives | Deterministic policy engine in the tool layer, plus policy text in context for explanation | Policy only in the prompt | Prompts are advisory; code is enforcement. The Air Canada tribunal case is the standard cautionary tale: the airline was held liable when its chatbot told a customer a bereavement refund could be claimed retroactively, contrary to its actual policy. Dentons 2024 |
| Single vs multi-agent | Single agent with ~10 tools | Specialist agents (billing, shipping) with a supervisor | Multi-agent adds hand-off failures and latency. Split only if the tool count or prompt size grows enough to hurt tool selection accuracy. |
| Model | Small model for routing and guards; mid-tier for the agent loop | Frontier everywhere | Tool-use reliability matters more than raw reasoning here; measure on the eval set before upgrading. |
| Fine-tuning | Not in v1 | Fine-tune on historic transcripts | Historic human transcripts encode old policies and human shortcuts. Revisit for routing and tone once there's logged, labeled agent traffic. |
Letting the model pass customer_id as a tool argument. A prompt-injected or confused model can then read someone else's orders. Bind identity on the server; tools take only the resource ID, and the tool layer checks that the resource belongs to the session's customer.
Evals
- Simulated-user, outcome-based evals. The τ-bench design is the right template: an LLM plays the customer with a hidden goal, the agent talks to it using real tools against a sandbox database, and success is checked by comparing the final database state to the expected state. Yao+ 2024 Outcome checks catch "said the right thing but did the wrong thing."
- Reliability, not just accuracy. τ-bench introduced pass^k (the probability of succeeding on all k independent trials) and found agents that look fine on a single run are much less consistent: retail pass^8 fell below 25% for the then state-of-the-art. Yao+ 2024 For support, where every customer is one trial, consistency is the metric that matters.
- Policy adversarial set. Customers asking for out-of-policy refunds, social engineering ("my manager said…"), injection inside order notes.
- LLM-as-judge on transcripts for tone, accuracy against policy, and escalation summary quality, calibrated against a human-labeled sample.
- Online: resolution rate, re-contact within 7 days, CSAT, escalation rate and reasons, refund dollars per conversation versus baseline, cost per conversation.
Failure modes
| Failure | Mitigation |
|---|---|
| Hallucinated policy ("you can return it after 90 days") | Ground answers in retrieved policy and cite the section; deterministic eligibility tool; judge-based spot checks for unsupported claims. |
| Wrong action (refunds the wrong item) | Confirmation step before mutations ("I'll refund the blue jacket, $59. Shall I proceed?"); idempotency keys; limits; reversible actions preferred. |
| Prompt injection via order notes or product reviews | Treat tool outputs as data; the tool layer enforces authorization regardless of what the model asks; no tool that sends data to arbitrary destinations. |
| Loops (agent keeps calling the same tool) | Step and token budgets; detect repeated identical calls; escalate on budget exhaustion. |
| Provider outage or latency spike | LLM gateway with fallback model; degrade to "FAQ plus escalate" mode. |
| Abuse and fraud (refund farming) | Fraud score in the eligibility tool; velocity limits per account; human approval above thresholds. |
How to extend
- Voice channel: reuse the same tool service and policy engine behind a voice front end (see case study 5). The tools are the durable asset; the conversational layer is swappable.
- Agent assist: run the same agent in "copilot" mode for human agents, suggesting replies and actions. That also generates labeled data (accepted, edited or rejected suggestions).
- Proactive support: trigger the agent from events (delayed shipment) with outbound messages.
- Distillation: once you have tens of thousands of successful, judge-verified trajectories, fine-tune a smaller model for the top intents.
The deep dive is almost always "how do you stop it doing something bad?" Strong answers layer the defenses: identity bound server-side; narrow tools; a deterministic policy engine; confirmation before mutations; amount limits and human approval; idempotency; full audit logs; and evals that check database outcomes, not just text. Mention that the company is legally responsible for what the bot says, so grounding and conservative escalation are product requirements, not nice-to-haves.
- τ-bench (Yao et al., 2024): simulated-user, database-state evaluation and the pass^k reliability metric.
- Anthropic: Writing effective tools for agents: tool naming, response shaping and tool evals.
- Anthropic customer support agent guide: a worked cost and prompt example.
Case study 2: Enterprise knowledge search (RAG) over Confluence, Drive and Slack with permissions
Prompt: "Employees waste hours finding information spread across Confluence, Google Drive and Slack. Design an assistant that answers questions over all of it, with citations, and never shows someone a document they can't access."
The interesting part isn't the LLM call. It's the ingestion pipeline, permission enforcement, freshness, and retrieval quality on messy enterprise data. Treat "never leaks a document" as a hard constraint, on par with correctness.
Requirements and clarifying questions
- Scale: how many employees and documents? Assume 20k employees; about 10M documents and messages, averaging 2k tokens, so roughly 20B tokens of raw text.
- Permission model: are source ACLs (Drive sharing, Confluence space and page restrictions, Slack channel membership) the source of truth? Yes. How quickly must a permission revocation take effect? Assume within 15 minutes for revocations; new content searchable within 1 hour.
- Answer style: a synthesized answer with citations, or ranked links? Both: answer plus sources, and a plain search mode.
- Latency: streaming answer, first token under 2 s, full answer under 8 s.
- Data residency and retention: must content stay in the company's cloud tenant? Assume yes; use a provider with zero data retention or a self-hosted model.
- Success metrics: answer correctness and citation faithfulness on a golden set; retrieval recall@k; zero permission leaks in red-team tests; online: weekly active usage, thumbs up rate, "answer found" rate, time saved in surveys.
Back-of-envelope
Assumptions: 20k users × 10 queries/day; chunks of about 500 tokens; 1024-dim embeddings; answer generation with a mid-tier model at $3 / $15 per M tokens.
| Quantity | Calculation | Result |
|---|---|---|
| Queries/day | 20k × 10 | 200k |
| QPS | 200k / 86,400, peak ×4 in business hours | about 2.3 avg, ~10 peak |
| Chunks | 20B tokens / 500 | 40M chunks |
| Vector storage | 40M × 4 KB (float32) / 1 KB (int8) / 128 B (binary) | 160 GB / 40 GB / 5 GB |
| Initial embedding cost | 20B tokens × $0.02–0.10 per M | $400–2,000 one-time |
| Optional contextual chunk headers | Anthropic estimated $1.02 per M document tokens with prompt caching Anthropic 2024; × 20B | ≈ $20k one-time (prices have fallen since, so likely less now) |
| Tokens per answer | 10 chunks × 500 + instructions and history ~2k → ~7–8k in; ~400 out | |
| Cost per answer | 8k × $3/M + 400 × $15/M | ≈ $0.03 |
| Daily generation cost | 200k × $0.03 | ≈ $6k/day (≈ $2k with a small model for easy queries) |
| Incremental ingestion | assume 1% of the corpus changes per day: 200M tokens | small: tens of dollars of embeddings per day |
Takeaway: generation dominates cost and retrieval quality dominates correctness; storage is cheap. 40M vectors fits comfortably in a single sharded vector index, or in Postgres with pgvector at the low end, and comfortably in a dedicated vector DB or a search engine with vector support (OpenSearch, Elasticsearch, Vespa).
Architecture
Component walkthrough
- Connectors. One per source, using webhooks or change feeds where available and periodic polling as a backstop. Each emits documents plus their ACLs and stable IDs. Commercial products in this space describe the same pattern: connectors mirror source permissions and resolve identities across systems so that if a user can't access an item in the source, they shouldn't see it in search. Glean docs
- Identity sync. Maps each source's users and groups to a canonical principal (the SSO identity plus its group memberships). Nested groups get expanded at sync time or at query time; expanding at query time keeps the index small but costs latency.
- Parse and chunk. Structure-aware chunking: split Confluence pages by headings, keep tables intact, convert Slack threads into a single "conversation" unit with the parent message as context, and drop bot noise. Each chunk carries metadata: source, title, author, last modified, URL, and a
acl_principalslist. - Enrichment. Prepending a short, model-generated context header to each chunk ("This section is from the 2026 travel policy, section on international flights…") helps both BM25 and embeddings. In Anthropic's experiments, contextual embeddings cut retrieval failures by 35%, adding contextual BM25 got to 49%, and adding reranking reached 67%. Anthropic 2024
- Hybrid index. BM25 catches exact terms (project code names, error codes, people's names) that embeddings blur; dense vectors catch paraphrases. Merge them with reciprocal rank fusion, \( \text{score}(d) = \sum_r \frac{1}{k + \text{rank}_r(d)} \) with \( k \approx 60 \), which needs no score calibration between the two retrievers.
- Query planner. A small model rewrites conversational follow-ups into standalone queries, extracts filters ("in the Payments space", "last quarter"), and decides whether the question needs retrieval at all, or a structured source (the HR system for "how many PTO days do I have").
- Retrieve, filter, rerank. Retrieve about 100 candidates with the ACL filter applied inside the index query (the chunk's principals intersect the user's), then a cross-encoder reranks to about 10.
- Live ACL re-check. For the final handful of documents, re-check access against the source API (cached for a few minutes). This covers the window between a revocation and the next sync.
- Answer with citations. The prompt requires every claim to cite a chunk ID, and a post-check verifies each citation actually supports its sentence (string overlap or a small NLI/judge model). Unsupported sentences are removed or flagged. If retrieval found nothing good, the model says so rather than answering from parametric memory.
Key design decisions and alternatives
| Decision | Options | Recommendation and reasoning |
|---|---|---|
| Permission enforcement point | (a) Post-filter results after retrieval; (b) pre-filter inside the index query; (c) separate index per user or group | (b) plus a live re-check. Post-filtering can return zero results when the top-k is mostly forbidden, and leaks via timing or counts. Per-user indexes don't scale with 20k users and overlapping groups. |
| Retrieval | Dense only; BM25 only; hybrid | Hybrid plus a reranker. Enterprise queries are full of rare identifiers. |
| Agentic vs single-shot retrieval | One retrieval call vs an agent that searches iteratively | Single-shot for most questions (fast, cheap). Escalate to an agentic multi-search mode for complex "compare / summarize across" questions, with a visible "researching…" state. |
| Long context instead of RAG | Stuff whole documents into a 1M-token window | Useful for "summarize this doc" on one or a few docs. Not a substitute at 20B tokens, and models use the middle of long contexts less reliably. Liu+ 2023 |
| Fine-tune on the corpus | Teach the model company facts | No. Facts change daily, fine-tuning can't enforce per-user permissions, and it can't cite. Retrieval is the right tool for knowledge; fine-tuning is for behavior and format. |
| Slack handling | Index each message; index threads; index channel-day windows | Threads as units, with channel name and date. Single messages lack context; day windows mix topics. |
Embedding documents with a service account and enforcing permissions "in the prompt" ("only use documents the user can see"). The model has no idea what the user can see, and anything in the context window can end up in the answer. Permissions must be enforced before content reaches the model.
Evals
- Retrieval evals (cheap, fast, run on every change): a few hundred real questions with labeled relevant documents; measure recall@10, recall@100 and MRR. Most RAG quality problems are retrieval problems, so measure them separately from generation.
- Answer evals: LLM judge with a rubric for correctness against a reference answer, citation faithfulness (every claim supported by its cited chunk), and appropriate abstention ("I couldn't find this").
- Permission red-team suite: synthetic users in specific groups asking about documents they shouldn't see, including indirect asks ("summarize the layoffs plan"). The pass criterion is zero leaks, run in CI.
- Freshness tests: update or revoke a document and measure time until the index reflects it.
- Online: thumbs, citation click-through, query reformulation rate (a proxy for a bad first answer), zero-result rate.
Failure modes
| Failure | Mitigation |
|---|---|
| Stale or conflicting documents (three versions of the travel policy) | Boost recency and "official" spaces; dedupe near-duplicates; show dates in citations; tell the model to surface conflicts rather than pick silently. |
| Permission leak through ACL lag | Event-driven revocation sync; live re-check on final documents; short cache TTLs. |
| Over-shared content (a doc shared "anyone in the company" that shouldn't be) | Not technically a leak, but AI search makes it findable. Offer admins an over-sharing report and sensitive-label exclusions before launch. |
| Prompt injection in an indexed doc ("ignore previous instructions and…") | Treat retrieved text as data, delimit it, give the answering model no side-effecting tools, and sanitize rendered links and images (no auto-loading images that could exfiltrate data through URLs). |
| Hallucinated answer when retrieval misses | Relevance threshold on reranker scores; abstain below it; citation verification. |
| Connector rate limits and API changes | Backoff, per-source queues, health dashboards, and alerting on ingestion lag. |
How to extend
- Actions: "file a Jira ticket from this thread." This turns a read-only system into one with side effects, which re-opens the injection threat model; add confirmations.
- Structured sources: text-to-SQL over the data warehouse, with row-level security enforced by the database, not by the model.
- Expose as an MCP server so other agents (coding agent, support agent) can call "company search" with the user's identity.
- Personalization: boost documents from the user's team and recent collaborators.
Expect a deep dive on permissions: "a user loses access to a doc; walk me through what happens." A strong answer covers the revocation event from the source, the identity and ACL sync, updating the chunk's principals in the index, the live re-check as a safety net, and what the worst-case window is. Also expect "how do you know retrieval is the problem and not generation?", which you answer with separate retrieval metrics and by looking at whether the right chunk was in the context for the failed answers.
- Anthropic: Contextual Retrieval: contextual chunk headers, hybrid BM25 plus embeddings, reranking, with measured gains.
- Lost in the Middle (Liu et al., 2023): why "just use long context" isn't free.
- Glean: How connectors power the Glean experience: a production description of permission mirroring and identity resolution.
Case study 3: Autonomous coding agent (issue → pull request)
Prompt: "Design a system where engineers assign a GitHub issue to an AI agent, and it comes back with a pull request that passes CI."
Coding is the domain where agents work best today, for one structural reason: there's a cheap, automatic verifier. Tests, type checkers, linters and builds give the agent a feedback signal it can iterate against, and give you an eval. Anthropic's coding-agent guidance boils down to "give the agent a check it can run," because without one, "looks done" is the only stopping signal. Claude Code docs
Requirements and clarifying questions
- Scope of tasks: bug fixes and small features (under ~300 lines changed) in existing repos, or greenfield? Assume the former.
- Repos: how many, which languages, how big? Assume 500 repos, mostly TypeScript, Python and Go, up to a few million lines.
- Autonomy: does it open PRs for human review (yes), or merge itself (no, never in v1)?
- Environment: can the agent build and run tests? It must be able to, or quality collapses. Do repos have reproducible dev environments (devcontainers, Dockerfiles)?
- Secrets and network: what may the sandbox reach? Package registries yes; production systems and secrets no.
- Volume: assume 2,000 tasks/day across the company.
- Success metrics: PR merge rate (merged with no or minor edits), CI pass rate at first PR, reviewer time per PR, time from assignment to PR, cost per merged PR, and revert rate after merge.
Back-of-envelope
Assumptions: a typical task takes about 50 agent steps (read files, grep, edit, run tests); the context grows from ~10k to ~80k tokens, averaging ~40k input tokens per step; ~400 output tokens per step; a mid-tier model at $3 / $15 per M; prompt caching on the growing prefix.
| Quantity | Calculation | Result |
|---|---|---|
| Input tokens per task | 50 steps × 40k | ~2M |
| Output tokens per task | 50 × 400 | ~20k |
| Uncached cost per task | 2M × $3/M + 20k × $15/M | ≈ $6.30 |
| Cached cost per task | ~90% of input is a cache hit (the prefix from previous steps): 1.8M × $0.30/M ≈ $0.54; ~0.2M new tokens written at 1.25× ≈ $0.75; output $0.30 | ≈ $1.60 |
| Daily model spend | 2,000 × $1.60–6.30 | ≈ $3k–13k/day |
| Wall-clock per task | 50 steps × (~3 s model + ~3 s tool/test) | ~5 min model time, often 10–30 min with test suites |
| Concurrent sandboxes | 2,000/day × 20 min ÷ 1,440 min/day, peak ×3 | ~28 avg, ~100 peak |
| Sandbox compute | 2 vCPU / 4 GB for 20 min at ~$0.05–0.10 per vCPU-hour | a few cents per task: negligible next to tokens |
Two important observations. First, caching is the difference between viable and not: agent loops re-send the whole growing transcript every step, so the uncached cost grows roughly quadratically with step count. Second, the right denominator is cost per merged PR, not cost per task: if only 40% of PRs merge, the effective cost is 2.5× higher, and a stronger, pricier model with a higher merge rate can be cheaper overall.
$$ \text{cost per merged PR} = \frac{\text{cost per attempt}}{\text{merge rate}} + \text{reviewer time} \times \text{engineer hourly cost} $$Reviewer time usually dominates: 15 minutes of an engineer's time is worth far more than the tokens. So optimize for PRs that are easy to review (small, well-described, with tests), not just for raw success rate.
Architecture
Component walkthrough
- Task intake. Normalize the issue into a task spec: description, repo, base branch, linked files, acceptance criteria if present. If the issue is too vague, the agent's first move should be to ask a clarifying question on the issue, not guess. A cheap triage model can score "is this well specified and small enough" and decline the rest.
- Sandbox. Each task gets an isolated, ephemeral environment: a microVM (Firecracker-style) or a hardened container (gVisor), with the repo checked out at a pinned SHA, dependencies pre-installed from a cached image, and a network egress allowlist. The OpenHands platform is a good public reference: agents act by writing code, using a shell and browsing inside sandboxed environments. Wang+ 2024 Anthropic describes the same principle for Claude Code: filesystem and network isolation enforced by OS primitives, which in their testing reduced permission prompts by 84%. Anthropic 2025
- Agent-computer interface. The SWE-agent paper's key finding is that the interface matters as much as the model: purpose-built commands for viewing files in windows, searching, and editing with immediate lint feedback beat dumping raw shell output on the model. Yang+ 2024 Practical rules: edits as diffs or string replacements, not whole-file rewrites; truncate long outputs with a pointer to see more; show line numbers.
- Repo context. A conventions file at the repo root (build and test commands, style rules, gotchas) that the agent reads first; a symbol index or LSP for "go to definition"; grep for exact strings. Embedding search over code helps for "where is the auth logic" questions, but most agents lean on grep and file navigation, which are exact and cheap.
- Orchestrator. A durable job (Temporal, a workflow engine, or a queue plus checkpointing) because tasks run for many minutes and must survive restarts. It enforces budgets (max steps, tokens, wall clock), compacts context when it nears the limit (summarize old steps, keep the plan and modified-file list), and can spawn sub-agents for isolated exploration so the main context stays clean. Anthropic 2025
- Verifier. Runs the full CI-equivalent suite independently of the agent's own claims. Optionally a reviewer agent in a fresh context looks only at the diff and the issue, which avoids the author's bias. A failed verification feeds back into the loop for a bounded number of retries.
- PR creation. A draft PR with a summary of the change, why, which tests were added or run (with output), and any uncertainty. Human merge only.
Key design decisions and alternatives
| Decision | Options | Recommendation and reasoning |
|---|---|---|
| Agent vs fixed pipeline | Agentless-style localize → repair → validate vs open-ended agent | Agent with a suggested plan structure. The fixed pipeline is a strong, cheap baseline (Agentless: 32% on SWE-bench Lite at $0.70/issue in 2024 Xia+ 2024), but real repos need running builds, reading errors and adapting. |
| Sampling strategy | One attempt vs N parallel attempts with selection | Start with one. For hard tasks, run 2–4 attempts in parallel sandboxes and pick the one that passes tests with the smallest diff. This trades cost for success rate, and test-based selection is reliable. |
| Isolation | Docker container; gVisor; microVM | MicroVM or gVisor for untrusted, model-generated code. Plain containers share the host kernel. |
| Network | Open; allowlist; none | Allowlist (package mirrors, internal docs). Open network plus repo secrets plus untrusted issue text is the lethal trifecta. Willison 2025 |
| Credentials | Developer's token vs scoped app token | A GitHub App token scoped to one repo, able to push to agent branches only. No deploy or production credentials in the sandbox at all. |
| Model | Frontier vs mid-tier | Measure cost per merged PR. Coding is where frontier models most often pay for themselves. |
Evals
- Public benchmark as a sanity check: SWE-bench consists of real GitHub issues from Python repos, graded by running the repo's tests (2,294 instances in the original set). Jimenez+ 2023 The human-validated "Verified" subset and other variants (Lite, Multilingual, Multimodal) are on the official leaderboard. swebench.com
- Internal benchmark, which is what actually matters: mine your own merged PRs that fixed issues and had tests. Reset the repo to the parent commit, give the agent the issue, and grade by running the PR's tests (fail-to-pass and pass-to-pass). A few hundred of these from your own repos predicts production far better than public numbers.
- Process metrics: steps, tokens and cost per task; tool error rate; how often the agent ran tests before finishing.
- Human-judged quality on a sample: would you merge this? Is the diff minimal and idiomatic? Did it add tests?
- Online: merge rate, reviewer edits, time to merge, revert rate, developer satisfaction.
Top SWE-bench Verified scores rose from the low teens (SWE-agent reported 12.5% on full SWE-bench in 2024 Yang+ 2024) to well above 70% by 2025–2026, and the community has raised concerns about saturation and contamination of the public set. Check the current leaderboard and newer benchmarks before quoting numbers, and in interviews, emphasize internal evals over public ones.
Failure modes
| Failure | Mitigation |
|---|---|
| Gaming the tests (deleting failing tests, special-casing inputs, weakening asserts) | Verifier flags test-file deletions and modified asserts; reviewer agent instructed to look for this; protected test directories; hidden held-out tests in evals. |
| Wandering (reads 200 files, never converges) | Step budget; plan-first prompting; sub-agents for exploration; stop and report when the budget runs out. |
| Context overflow on large repos | Compaction; truncated tool outputs; just-in-time file reads rather than pre-loading. |
| Prompt injection through issue text, code comments or dependency READMEs | No secrets in the sandbox; egress allowlist; scoped tokens; human merge gate. |
| Flaky tests misleading the agent | Run the baseline suite before changes to record pre-existing failures; rerun failures once. |
| Huge, unreviewable PRs | Diff-size limit; ask the agent to split; decline tasks that triage scores as too big. |
How to extend
- CI-failure fixer and dependency upgrader: high-volume, well-verified task types that are a natural next step.
- Parallel fan-out for migrations: split a large migration into per-file tasks, run many agents, and merge in batches.
- Learning from review: feed reviewer comments into the repo's conventions file, so the next PR avoids the same mistake.
- Interactive mode: the same backend serving a developer in their IDE, with the human as the verifier.
Interviewers probe the sandbox ("what stops it from curl-ing your secrets to an attacker?") and the verification story ("how do you know the PR is correct?"). A strong answer separates the agent's own checks from an independent verifier, mentions test gaming as a known failure, and frames the metric as cost per merged PR including reviewer time. Bonus: explain why caching makes or breaks the economics of long agent loops.
- SWE-agent (Yang et al., 2024): why the agent-computer interface matters.
- Agentless (Xia et al., 2024): the strong fixed-pipeline baseline.
- Anthropic: Beyond permission prompts (sandboxing): filesystem and network isolation for a coding agent.
- Claude Code best practices: verification, context management, subagents and fan-out in practice.
Case study 4: Deep-research agent (orchestrator with parallel web-search subagents and citations)
Prompt: "Design a 'deep research' feature: the user asks a complex question, and the system spends several minutes searching the web and produces a long, well-cited report."
This is the canonical case for multi-agent systems, and there's an unusually detailed public write-up: Anthropic's description of how they built Claude's Research feature. Anthropic 2025 OpenAI's deep research, launched in early 2025, was described as an agent powered by a version of o3 optimized for browsing, searching and analyzing large amounts of web content. OpenAI 2025 Use these as reference points, but design from first principles.
Requirements and clarifying questions
- Query types: breadth-first ("compare the 10 leading vector databases on X, Y, Z") vs depth-first ("why did this drug fail its phase III trial")? Both, which is why the planner matters.
- Latency tolerance: users accept minutes for a report; assume a 3–15 minute target with visible progress.
- Output: a structured report of 1,500–5,000 words with inline citations to specific sources.
- Sources: open web only, or also the user's documents and internal tools? Start with web, design for pluggable sources.
- Volume: assume 5,000 reports/day.
- Success metrics: factual accuracy, citation accuracy (does the cited page actually support the claim), completeness against what an expert would cover, source quality, and efficiency (tokens and tool calls). These are close to the rubric Anthropic described for its judge. Anthropic 2025
Back-of-envelope
Assumptions: a lead agent plus 5 subagents on average; each subagent makes about 20 tool calls (searches and page fetches) and averages about 30k tokens of context per step; a final citation pass; mid-tier model at $3 / $15 per M; web search at $10 per 1,000 searches (Anthropic's listed price for its server-side search tool). Anthropic pricing
| Component | Input tokens | Output tokens |
|---|---|---|
| Lead agent: plan, dispatch, read results, synthesize | ~200k | ~15k (the report) |
| 5 subagents × 20 steps × ~30k context | ~3M | 5 × 10k = 50k |
| Citation pass | ~100k | ~5k |
| Total | ~3.3M | ~70k |
- Uncached: 3.3M × $3/M + 70k × $15/M ≈ $9.90 + $1.05 ≈ $11 per report. With caching inside each subagent loop, roughly $4–6.
- Searches: ~75 per report × $0.01 ≈ $0.75.
- 5,000 reports/day → roughly $25k–55k/day. This is why deep research is often rate-limited per user or tied to higher plans.
- Sanity check against the published ratio: Anthropic reported agents use about 4× the tokens of a chat interaction and multi-agent systems about 15×. Anthropic 2025 If a chat turn is a few thousand tokens, 15× is tens of thousands per turn, and a long research run covers many turns, so millions of tokens is plausible.
- Concurrency: 5,000/day × 10 min ÷ 1,440 ≈ 35 concurrent research jobs on average, each fanning out to ~5 subagents, so ~200 concurrent agent loops at peak. Plan rate limits with your model provider accordingly.
Architecture
Component walkthrough
- Clarification. Before spending dollars, ask one or two clarifying questions if the query is ambiguous ("for which market? as of when?"). Cheap and high leverage.
- Lead agent (orchestrator). Writes a research plan, saves it to memory (so it survives context truncation), and spawns subagents with explicit task descriptions: objective, output format, which tools and sources to prefer, and boundaries so subagents don't duplicate each other. Anthropic found vague delegation caused duplicated work and gaps, and that teaching the lead to scale effort to query complexity (one agent for simple fact-finding, many for broad comparisons) mattered a lot. Anthropic 2025
- Subagents (workers). Each runs its own search loop: broad queries first, then narrower ones; fetch and read pages; judge source quality; return a condensed summary with source URLs and key quotes. The main benefit is context isolation: each subagent burns through 100k+ tokens of web pages, but returns only 1–2k tokens to the lead.
- Parallelism. Anthropic reported that having the lead spawn 3–5 subagents in parallel, and subagents call tools in parallel, cut research time by up to 90% for complex queries. Anthropic 2025
- Iteration. After a round, the lead evaluates coverage against the plan and either spawns another round for gaps or proceeds to writing.
- Citation agent. A separate pass maps each claim in the draft to the specific source passage that supports it. Separating "write" from "cite" makes it easier to verify and to catch unsupported claims.
- Source store. Fetched pages are cached and content-addressed, so the citation agent and the UI can show exactly what was read, and repeated fetches are free.
- Durable execution. Long, stateful runs need resumability: if a step fails, resume from a checkpoint instead of restarting a 10-minute job. Anthropic also mentions deploying updates with "rainbow deployments" so running agents aren't broken by a mid-run code change. Anthropic 2025
Key design decisions and alternatives
| Decision | Options | Recommendation and reasoning |
|---|---|---|
| Single agent vs multi-agent | One long-context agent doing everything sequentially vs orchestrator plus parallel workers | Multi-agent fits here because research is breadth-first and parallelizable with mostly independent subtasks. Anthropic's multi-agent setup (Opus 4 lead, Sonnet 4 subagents) beat single-agent Opus 4 by 90.2% on their internal research eval, and token usage alone explained 80% of the variance on BrowseComp. Anthropic 2025 |
| When not to use multi-agent | Tasks where steps depend tightly on each other, or where all agents need shared context, such as most coding. Cognition argued that parallel subagents make conflicting implicit decisions unless they share full context and traces. Yan 2025 The two posts aren't really contradictory: they describe different task shapes. | |
| Model mix | Same model everywhere vs strong lead plus cheaper workers | A strong lead (planning and synthesis are the hard parts) with mid-tier workers. Measure whether upgrading workers moves the eval. |
| Search provider | Commercial search API, provider-native search tool, own crawler | Commercial or provider-native search for v1. Your own index only for a vertical (legal, academic) where you control the corpus. |
| Report generation | Single pass vs outline then section-by-section | Outline first, then sections, for long reports: it keeps each generation call focused and lets you stream progress. |
Multi-agent systems are mostly a way to spend more tokens productively in parallel. If the task's quality scales with how much material you read and the subtasks are separable, more parallel context windows win. If the task needs one coherent line of reasoning, they mostly add coordination failures.
Evals
- Start small, immediately. Anthropic started with about 20 queries representing real usage, because early changes have large effects that a small set can detect. Anthropic 2025 Don't wait to build a 1,000-item eval before iterating.
- LLM judge with a rubric: factual accuracy, citation accuracy, completeness, source quality, tool efficiency. A single judge call scoring 0–1 plus pass/fail was reported to be more consistent than multiple specialized judges. Anthropic 2025
- Short-answer benchmarks for search ability: BrowseComp (1,266 hard-to-find but easily verified questions) tests persistent browsing. Wei+ 2025
- Citation verification at scale: automatically check that each cited URL was actually fetched and that the quoted passage appears on it.
- Human review: Anthropic's human testers found the agent preferred SEO-optimized content farms over authoritative sources, which automated evals had missed. Anthropic 2025
Failure modes
| Failure | Mitigation |
|---|---|
| Fabricated or mismatched citations | Separate citation pass; only allow citing fetched sources from the source store; automated quote-presence checks. |
| Low-quality sources (SEO spam, AI-generated content farms) | Source-quality heuristics in the subagent prompt; domain allow/deny lists; prefer primary sources. |
| Over-spawning (50 subagents for a simple question) | Explicit effort-scaling rules in the lead prompt; hard caps on subagents and tool calls; budget per report. |
| Duplicated or overlapping subagent work | Clear task boundaries in delegation; lead dedupes findings. |
| Prompt injection from web pages | Web content is untrusted; subagents have no side-effecting tools; never combine web browsing with access to private user data and an outbound channel without strict controls. |
| Errors compounding over a long run | Checkpoints and resume; let the agent know when a tool failed so it can adapt; tracing of every decision for debugging. |
How to extend
- Private sources: add enterprise search (case study 2) as a tool, with the user's permissions. Now the trifecta applies: web content plus private data. Restrict outbound actions and render links safely.
- Async and scheduled research: "monitor this topic weekly", with diffs against the previous report.
- Structured outputs: tables and charts from extracted data via code execution.
- Training: RL on research trajectories with verifiable rewards (short-answer questions) is how frontier labs improved browsing agents; out of scope for most product teams.
Expect "why multi-agent instead of one agent with a big context?" The strong answer: context isolation (each worker can read 100k+ tokens and return a summary), parallelism (wall-clock time), and the evidence that more tokens spent productively correlates with quality on this task type. Then volunteer the costs: roughly 15× chat tokens, coordination failures, harder debugging, and that it's a poor fit for tightly coupled tasks. Showing both sides is the signal.
- Anthropic: How we built our multi-agent research system: the best public write-up of a production deep-research architecture, prompts and evals.
- Cognition: Don't build multi-agents: the counter-argument about shared context.
- BrowseComp (Wei et al., 2025): a benchmark for persistent web browsing.
Case study 5: Real-time voice agent
Prompt: "Design a voice agent that answers inbound phone calls for a chain of clinics: it books, moves and cancels appointments, and answers basic questions."
Everything from case study 1 applies (tools, policy, escalation), but voice adds a hard real-time constraint. Humans expect a reply within a few hundred milliseconds of finishing a sentence; a 3-second pause feels broken. The design is dominated by the latency budget, turn-taking, and telephony plumbing.
Requirements and clarifying questions
- Channel: PSTN phone calls (inbound and possibly outbound), or in-app WebRTC? Phone, which means 8 kHz narrowband audio and extra network hops.
- Latency target: voice-to-voice (user stops talking → agent audio starts) under ~1.2 s at p50, under 2 s at p95.
- Languages and accents, and handling of names, dates of birth and phone numbers (where speech recognition is error-prone).
- Compliance: healthcare means HIPAA-style requirements: business associate agreements with every vendor, call recording consent, PHI handling, and transcripts for audit.
- Volume: assume 100k calls/day, averaging 4 minutes.
- Hand-off: warm transfer to front-desk staff with a summary.
- Success metrics: task completion rate (booking made correctly), transfer rate, average handle time, voice-to-voice latency percentiles, interruption-handling errors (talking over the caller, cutting them off), caller satisfaction, cost per call.
Two architectures: cascaded pipeline vs speech-to-speech
Cascaded: STT → LLM → TTS
- Streaming speech-to-text produces partial and final transcripts.
- A text LLM with tools generates a streamed reply.
- Streaming text-to-speech starts speaking from the first sentence fragment.
- Pros: pick the best component for each stage; full text transcript for compliance, evals and debugging; mature tool calling; easy context control; cheaper.
- Cons: more moving parts; latency adds up across stages; loses prosody (tone, hesitation) from the input.
Speech-to-speech (native audio model)
- One model takes audio in and produces audio out (e.g. OpenAI's Realtime API over WebRTC or WebSocket). OpenAI docs
- Pros: potentially lower latency; more natural prosody and emotion; hears tone and non-verbal cues.
- Cons: one practitioner guide estimates it at 3–5× the cost of a cascaded pipeline, with less flexible context management and no reliable transcript for compliance Voice AI primer 2026; harder to eval; tool-use reliability historically behind text models.
For a clinic-booking agent with compliance needs and heavy tool use, the cascaded pipeline is the defensible default in 2026. Speech-to-speech fits consumer companionship or language-practice products where naturalness beats auditability. Say that this is a fast-moving trade-off.
Speech-to-speech models improved rapidly in 2025–2026, and the cost and tool-calling gap has been narrowing. Hybrid designs (audio-native understanding with text-based reasoning and tools) are also appearing. Check current model capabilities and pricing before committing to either side in a real design.
Latency budget
The most useful public breakdown comes from a widely used practitioner primer (originally written for the AI Engineering Summit in February 2025, updated in 2026), which totals roughly 1.3 s for a typical cascaded pipeline and recommends about 1.5 s as an achievable target (with well-optimized stacks reaching around 500 ms). Voice AI primer 2026
| Stage | Typical (primer) | How to shrink it |
|---|---|---|
| Mic capture and client buffering | ~40 ms | Small frames (20 ms); hardware echo cancellation. |
| Network, encoding and decoding (both directions) | ~140 ms total | WebRTC over UDP rather than WebSocket over TCP; regional edge servers; co-locate STT, LLM and TTS. |
| Transcription plus endpointing (deciding the user is done) | ~300 ms | Streaming STT; a smart turn-detection model instead of a long silence timeout. |
| LLM time to first token | ~650 ms | Smaller or faster model; short prompts; prompt caching; avoid reasoning modes on the hot path; keep tool calls off the critical path when possible. |
| TTS time to first byte | ~120 ms | Streaming TTS fed sentence by sentence. |
| Total voice-to-voice | ~1.3 s |
Telephony adds more: PSTN carriers and the media relay add their own buffering, so budget a few hundred extra milliseconds for phone calls compared to WebRTC (I don't have a verified source for a specific figure; measure it). Tool calls are the other big latency risk: a 1.5 s appointment-system lookup in the middle of a turn blows the budget, so play a short filler ("Let me check that for you…") while it runs.
$$ T_{\text{v2v}} = T_{\text{net}} + T_{\text{endpoint}} + T_{\text{STT,final}} + \text{TTFT}_{\text{LLM}} + T_{\text{tool}}\ (\text{if any}) + \text{TTFB}_{\text{TTS}} $$Turn detection and barge-in
Turn-taking is the hardest part to get right and the part users notice most.
- Endpointing (when has the user finished?). The naive approach is voice activity detection (VAD) plus a silence timeout. OpenAI's Realtime API defaults to 500 ms of silence in its
server_vadmode and also offers asemantic_vadmode that uses a classifier over the words spoken to decide whether the user is done. OpenAI docs Silence alone fails both ways: people pause mid-sentence ("my date of birth is… uh… March 3rd"), and a long timeout adds dead air to every turn. - Turn-detection models. Open models now do this better. Pipecat's Smart Turn is an ~8M-parameter model that works directly on audio (16 kHz, up to 8 s), using prosody as well as content, with inference around 10–100 ms on CPU. Pipecat Smart Turn LiveKit shipped a transcript-based turn detector built on a small Qwen model and has since moved toward an audio-based one. LiveKit docs The emerging pattern combines a short VAD trigger, an audio turn model, and sometimes the LLM itself emitting a "complete / incomplete" token. Voice AI primer 2026
- Barge-in (the user interrupts the agent). When VAD detects user speech while the agent is talking, immediately stop TTS playback, flush buffered audio (on Twilio Media Streams, send a
clearmessage to drop queued audio Twilio docs), cancel the in-flight LLM generation, and start listening. - Context truncation after barge-in. The LLM's transcript says it said a full paragraph, but the user only heard the first sentence. Use word-level timestamps from TTS (or playback "mark" events) to truncate the assistant message to what was actually played, so the model doesn't assume the user heard information they didn't. Voice AI primer 2026
- Backchannels and false barge-ins. "Mm-hm" or a cough shouldn't stop the agent. Require a minimum speech duration or a short transcript before treating it as an interruption.
Back-of-envelope
Assumptions: 100k calls/day × 4 min = 400k minutes/day; about 6 conversational turns per minute (both sides); the per-minute component prices below are rough ranges, flagged as such.
| Quantity | Calculation | Result |
|---|---|---|
| Average concurrent calls | Little's law: (100k / 86,400 s) × 240 s | ≈ 280 |
| Peak concurrent calls | clinic hours concentrate traffic; ×3–4 | ≈ 1,000 |
| LLM tokens per minute | ~3 agent turns/min × ~4k-token context (mostly cached) + ~60 output tokens | ~12k in, ~200 out |
| LLM cost per minute | with caching on a small or mid model; the primer estimates roughly $0.002–0.02 per 3-minute call for common models Voice AI primer 2026 | ~$0.001–0.01 |
| STT per minute | streaming STT list prices | ~$0.004–0.01 |
| TTS per minute | primer: roughly $0.009–0.05 per minute at scale Voice AI primer 2026 | ~$0.01–0.05 |
| Telephony per minute | inbound PSTN plus media streaming | ~$0.005–0.02 |
| Total per minute | ~$0.03–0.09 | |
| Daily | 400k min × $0.03–0.09 | ≈ $12k–36k/day |
Notice that in voice, the LLM is often not the biggest line item: TTS and telephony are. A speech-to-speech model at several times the cost changes that picture, which is why the architecture choice is also a cost decision. The concurrency figure drives infrastructure: about 1,000 simultaneous WebSocket or WebRTC sessions, each holding audio buffers, VAD and turn models, and STT and TTS streams. Size media servers by concurrent sessions, not by QPS.
Architecture
Component walkthrough
- Telephony. A SIP trunk or a CPaaS provider answers the call and streams audio to your worker. Twilio's Media Streams, for example, forks call audio over a WebSocket as 8 kHz μ-law frames and accepts audio back the same way, with
markandclearmessages for playback tracking and interruption. Twilio docs At high volume, teams run their own SIP infrastructure to cut per-minute cost and latency. For in-app voice, use WebRTC end to end. - Voice agent worker. One stateful session per call, pinned to a region near the caller and the model endpoints. Frameworks such as Pipecat and LiveKit Agents provide this pipeline (transport, VAD, STT, LLM, TTS, interruption handling) so you don't hand-roll audio plumbing.
- STT. Streaming recognition with custom vocabulary (clinic names, doctor names, medication names). Confirm critical entities by reading them back ("That's March 3rd, 1985, correct?"), and spell out alphanumerics for confirmation.
- Dialogue LLM. A fast model with a short, cached system prompt, a small set of tools, and voice-specific instructions: short sentences, no markdown or lists, numbers written to be spoken, one question at a time. Stream tokens to TTS at sentence boundaries.
- TTS. Streaming synthesis with word-level timestamps; a consistent voice; pronunciation dictionaries for names.
- Tools. The same deterministic tool layer pattern as case study 1: search slots, book, reschedule, cancel, with identity verification (name plus date of birth plus phone match) enforced in code before any PHI is spoken.
- Observability. Per-stage latency for every turn (STT final, TTFT, TTFB), interruption events, and full recordings and transcripts (with consent) for replay-based evals.
Key design decisions and alternatives
| Decision | Chosen | Alternative and trade-off |
|---|---|---|
| Pipeline | Cascaded | Speech-to-speech: more natural, but costlier, harder to audit and eval. Revisit as models improve. |
| Endpointing | VAD + audio turn model | Fixed silence timeout: simpler, but either cuts people off or adds dead air. |
| LLM size | Fast, small-to-mid model on the hot path | Frontier model: better reasoning, but TTFT can blow the budget. Escalate hard cases to a slower path with a filler phrase, or transfer. |
| Hosting | Co-located regional workers, providers in the same region | Mixed regions: every cross-region hop adds tens of milliseconds per stage. |
| Speculative work | Start LLM generation on a confident partial transcript; discard if the user keeps talking | Wait for the final transcript: simpler, but slower. Speculation costs extra tokens. |
Evals
- Simulated callers: an LLM caller plus TTS with varied voices, accents, background noise and line quality, running scripted scenarios against the real pipeline. Score task completion by checking the scheduling system's final state.
- Component evals: STT word error rate on your own call audio (especially names and numbers), turn-detection precision and recall (cut-offs vs late responses), TTS pronunciation.
- Latency SLOs: p50 and p95 voice-to-voice per stage, tracked in production dashboards.
- Transcript-level LLM judge for policy adherence, tone, and whether identity was verified before PHI was disclosed.
- Online: completion and transfer rates, call abandonment, repeat calls, post-call surveys.
Failure modes
| Failure | Mitigation |
|---|---|
| Talking over the caller, or cutting them off mid-sentence | Better turn model; tune per use case (people reading out numbers pause a lot); allow quick recovery ("sorry, go ahead"). |
| Misheard entities (wrong date, wrong name) | Read-back confirmation for anything that will be written; DTMF keypad fallback for numbers. |
| Latency spikes from a provider | Per-stage timeouts; fallback providers through a gateway; filler phrases; graceful transfer. |
| Disclosing PHI to the wrong person | Identity verification enforced by tools before data access; tools refuse otherwise. |
| Agent doesn't know the user didn't hear something (after barge-in) | Truncate context with playback timestamps. |
| Robocall and abuse traffic | Rate limits per number, spam scoring, maximum call duration. |
How to extend
- Outbound calls (reminders, rescheduling after a cancellation), which bring consent and calling-regulation requirements.
- Multilingual with language detection at the start of the call.
- Hybrid audio understanding: use an audio-native model to detect frustration or confusion, and escalate.
- Self-hosted STT and TTS at high volume to cut per-minute cost and latency variance.
The question is almost always "walk me through the latency budget" followed by "what happens when the user interrupts?" A strong answer gives per-stage numbers, explains why endpointing is a hidden latency cost, describes streaming at every stage, and covers barge-in end to end, including truncating the context to what was actually heard. Bonus points for noting that LLM tokens are often not the dominant cost in voice, and for treating concurrency (not QPS) as the scaling unit.
- Voice AI & Voice Agents: An Illustrated Primer: the most complete practitioner guide; latency tables, turn detection, cost estimates (updated 2026).
- Pipecat Smart Turn: an open audio-based turn-detection model.
- OpenAI Realtime API: voice activity detection: server VAD vs semantic VAD in a speech-to-speech API.
- Twilio Media Streams messages: what the telephony side of a voice agent actually exchanges.
Case study 6: Personalized assistant with long-term memory
Prompt: "Design a consumer AI assistant that remembers users across sessions (their preferences, projects, people in their life) and uses that to personalize answers, while respecting privacy."
Every major assistant now ships some form of memory. ChatGPT, for example, distinguishes explicit "saved memories" from implicitly "referencing chat history", and lets users view, delete and turn off both. OpenAI 2024 The design challenge isn't storing text. It's deciding what is worth remembering, keeping it correct as facts change, retrieving the right memory at the right time without creeping users out, and giving users real control.
Requirements and clarifying questions
- What kinds of memory? Stable preferences ("vegetarian", "prefers concise answers"), facts about the user's life ("has a daughter named Maya", "training for a marathon in April"), ongoing projects, and episodic recall ("what was that book we discussed last month?").
- Explicit vs implicit: only remember when asked ("remember that…"), or infer automatically? Both, with implicit memories visible and editable.
- Scale: assume 1M daily active users, 10 messages each per day, about 3 sessions per user per day.
- Privacy and regulation: GDPR and CCPA rights (access, deletion), sensitive categories (health, religion, sexuality, political views), minors, and regional data residency.
- Latency: memory retrieval must add less than ~150 ms to the first token.
- Success metrics: memory precision (stored memories are true and useful), recall in context (the right memory shows up when relevant), inappropriate-use rate (memory surfaced where it feels intrusive), user edits and deletions of memories, retention and satisfaction for memory-on vs memory-off cohorts.
Memory types and where they live
| Type | Example | Storage | How it reaches the model |
|---|---|---|---|
| Working memory | The current conversation | Context window, compacted when long | Always in context |
| Profile / core facts | Name, language, dietary preference, tone preference | Small structured record per user (a few hundred tokens) | Always injected into the system prompt (and cacheable per user) |
| Semantic memories | "Training for the Berlin marathon in April 2027" | Memory store: text + embedding + metadata (source, time, confidence, sensitivity) | Retrieved per turn by relevance |
| Episodic memory | Summaries of past conversations | Session summaries, plus raw transcripts in cold storage | Retrieved on demand ("last month we talked about…") or via a search tool |
| Procedural | "When writing emails for me, sign off with 'Best, J'" | Instructions list | Injected as user-specific instructions |
This mirrors the operating-system analogy from MemGPT: a small, fast "main context" and larger external tiers that the model pages information in and out of, through function calls. Packer+ 2023
Back-of-envelope
Assumptions: 1M DAU × 10 messages = 10M messages/day; 3M sessions/day; on average 500 memories per long-term user after a year; mid-tier model at $3 / $15 for chat; small model at $1 / $5 (half price via batch) for memory extraction.
| Quantity | Calculation | Result |
|---|---|---|
| Chat QPS | 10M / 86,400, peak ×3 | ~115 avg, ~350 peak |
| Memory retrieval QPS | one retrieval per message | same: ~350 peak vector queries/s, each filtered to one user |
| Memory store size | (say) 5M registered users × 500 memories × ~2 KB (text + int8 vector + metadata) | ~5 TB, partitioned by user |
| Extra tokens per chat call from memory | profile ~300 + top-k memories ~700 | ~1k tokens → about $0.003 per message at $3/M |
| Daily added chat cost from memory | 10M × $0.003 | ≈ $30k/day (less with per-user profile caching) |
| Extraction cost | 3M sessions × ~3k tokens in, ~300 out, small model at batch price ($0.50 / $2.50 per M) | ≈ $4.5k + $2.3k ≈ $7k/day |
The interesting insight: injecting memory into every call costs more than extracting it, so keep the always-on profile compact and make retrieved memories earn their place (a relevance threshold, not a fixed top-k). Because each user's memory is small and queries are always scoped to one user, you don't need a huge global ANN index: partition by user, and a per-user brute-force search over 500 vectors takes microseconds.
Architecture
Component walkthrough
- Extraction. After a session (or every N turns), a small model reads the transcript and proposes candidate memories as structured objects:
{text, type, entities, valid_from, confidence, sensitivity, explicit}. Doing it asynchronously keeps it off the latency path and allows batch pricing. - Consolidation. The hard part. Each candidate is compared against the user's nearest existing memories, and an LLM decides add, update, delete or no-op. Mem0 popularized this extract-then-update pipeline, reporting large reductions in latency and token cost against full-context baselines on the LoCoMo benchmark (91% lower p95 latency and over 90% token savings, by their own measurements). Chhikara+ 2025 Treat vendor-authored benchmark numbers with appropriate caution.
- Temporal validity. Facts expire or change: "lives in Boston" becomes "moved to Seattle." Store
valid_fromandsuperseded_byinstead of overwriting, so the assistant can reason about time ("you used to live in Boston") and so a bad update can be rolled back. - Retrieval. At query time, a hybrid search within the user's partition, a relevance threshold, and a sensitivity filter (don't bring up a health memory in an unrelated work conversation). Optionally, give the model a
memory_searchtool and let it decide when to look things up, the "just-in-time" pattern. Anthropic 2025 Anthropic's API offers a client-side memory tool where the model reads and writes files under a/memoriespath that the application maps to its own storage. Anthropic docs - User controls. A memory page listing everything stored, with edit and delete; a "memory updated" indicator when something new is saved; a temporary or incognito mode; and a global off switch. These are product requirements, not polish.
- Deletion pipeline. Deleting a memory must also delete derived artifacts: embeddings, session summaries that mention it, caches, and (on a schedule) backups. Build this from day one; retrofitting it is painful.
Key design decisions and alternatives
| Decision | Options | Recommendation and reasoning |
|---|---|---|
| Just use long context | Stuff all past conversations into a 1M-token window | Too expensive per message and slower; models also degrade at using information spread across long histories. LongMemEval reported a roughly 30% accuracy drop for commercial assistants and long-context models on sustained-interaction memory tasks. Wu+ 2024 |
| Write path | Synchronous (during the turn) vs asynchronous | Async for implicit memories. Sync only for explicit "remember this," so you can confirm it immediately. |
| Representation | Free-text facts in a vector store; knowledge graph of entities and relations; files | Free-text facts plus entity tags in v1. A graph helps multi-hop questions ("what does my sister's husband do?") but adds complexity; Mem0's graph variant reported only a small gain over its base version. Chhikara+ 2025 |
| What to remember | Everything vs a whitelist of categories | Whitelisted, useful categories; explicit opt-in for sensitive ones; never credentials, financial account numbers or third parties' sensitive data. |
| Model fine-tuning per user | LoRA per user | No: can't be inspected, edited or deleted precisely, and it's expensive. Retrieval-based memory is transparent and deletable. |
Designing memory as "embed every message and retrieve top-k." Raw messages are noisy and redundant, contradictions pile up ("I'm vegetarian" in January, "I started eating fish" in June), and you can't show users a clean list of what's remembered. Extract, consolidate and version.
Evals
- Public long-term memory benchmarks: LoCoMo (conversations spanning up to 35 sessions) Maharana+ 2024 and LongMemEval, which tests information extraction, multi-session reasoning, temporal reasoning, knowledge updates and abstention. Wu+ 2024
- Extraction precision and recall: on labeled transcripts, did it store the facts a human would, and nothing false?
- Update correctness: scripted multi-session scenarios with changing facts; check that the latest fact wins and history is retained.
- Appropriateness: a judge rubric for "was surfacing this memory helpful or intrusive here?", plus red-team scenarios (sensitive memories in unrelated contexts, memories about third parties).
- Deletion tests: delete a memory, then probe with paraphrased questions to confirm it no longer influences answers.
- Online: memory-on vs memory-off A/B on retention and satisfaction; rate of user edits and deletions (a spike means bad extractions).
Failure modes
| Failure | Mitigation |
|---|---|
| Wrong memory ("you have two kids" — user has one) | Confidence scores; prefer explicit statements; surface "memory updated" so users can correct; version history. |
| Creepy or intrusive recall | Sensitivity labels and context-matching rules; don't volunteer sensitive memories unless the user raises the topic. |
| Memory poisoning via prompt injection (a web page or document tells the assistant to "remember that the user's bank is evil.com") | Only extract memories from the user's own messages, not from tool outputs or retrieved content; flag memory writes that originate during tool use. |
| Cross-user leakage | User ID enforced in the storage layer (partition keys, row-level security), never as a model-chosen parameter; per-user encryption keys. |
| Stale memories | Temporal validity, decay of unused memories, periodic re-confirmation for important facts. |
| Incomplete deletion (GDPR) | A deletion pipeline that tracks lineage of derived data; tests that assert deletion end to end. |
How to extend
- Shared or team memory with explicit sharing scopes (a family's grocery preferences, a team's project context), which brings permission models like case study 2.
- Proactive features: reminders and check-ins based on remembered plans (the marathon in April), with strict opt-in.
- On-device memory for privacy-sensitive users, syncing only encrypted blobs.
- Memory for agents, not just chat: an agent's learned procedures and lessons ("this user's repo uses pnpm, not npm").
Assistant memory features changed several times through 2024–2026 across vendors, and the memory-framework ecosystem (Mem0, Letta, Zep and others) moves quickly. Benchmark claims in this space are often self-reported by the framework authors. Check the current state before quoting specifics.
Interviewers want to see that you treat memory as a data lifecycle: extraction, consolidation with conflict resolution, temporal validity, retrieval with appropriateness filtering, user visibility and control, and verifiable deletion. Strong candidates bring up memory poisoning through prompt injection and the cost of injecting memory on every turn, and explain why per-user fine-tuning is the wrong tool.
- MemGPT (Packer et al., 2023): memory tiers and self-directed paging via function calls.
- Mem0 (Chhikara et al., 2025): the extract → consolidate (add/update/delete) pipeline with benchmarks.
- LongMemEval (Wu et al., 2024): what "good long-term memory" means, broken into testable abilities.
- Anthropic memory tool docs: a concrete client-side memory interface, including security considerations.
Three more prompts, sketched
These come up often enough that you should have a two-minute answer ready. Each one lists the clarifying questions that matter, the core architecture, the key numbers, and the deep-dive topics interviewers usually pick.
Sketch A: An internal LLM gateway / router platform
Prompt: "Fifty teams at our company call various LLM providers directly. Design a central platform for it."
- Clarify: which providers and self-hosted models; whether the goal is cost control, reliability, compliance, or developer velocity (usually all four); streaming support; latency overhead budget (aim for under ~20 ms added at p50).
- Core components: an OpenAI-compatible (or provider-neutral) API; authentication with per-team virtual keys; per-key budgets, rate limits and quotas; a provider abstraction layer; retries with backoff and fallback chains across providers and models, including fallbacks for context-window and content-policy errors (open-source gateways like LiteLLM implement exactly these patterns LiteLLM docs); request and response logging with PII redaction; cost attribution per team; guardrail hooks (PII, injection, moderation); and caching (exact-match and, carefully, semantic).
- Routing: static (team X uses model Y), rule-based (by task tag, prompt length or language), or learned. RouteLLM trained routers on preference data to choose between a strong and a weak model per query and reported cost reductions of over 2× in some settings without quality loss. Ong+ 2024 Learned routing needs per-use-case evals to prove it isn't hurting quality.
- Numbers: the gateway itself is a stateless proxy, so it scales horizontally; the hard limits are provider rate limits (tokens per minute) shared across teams. Plan capacity in tokens/minute, not requests, and do admission control by priority (interactive traffic over batch).
- Deep dives: streaming through the proxy (server-sent events, partial failures mid-stream: you can't transparently retry after tokens have reached the client); semantic-cache risk (serving the wrong answer to a similar-looking prompt, and cross-tenant leakage, so scope caches per tenant); model version pinning and deprecation migrations; data residency routing.
Sketch B: Document extraction pipeline for invoices at scale
Prompt: "We receive 2 million invoices a month as PDFs and scans from thousands of suppliers. Extract structured fields into our ERP."
- Clarify: the field list (supplier, invoice number, dates, line items, totals, tax, currency); the accuracy bar per field (totals and bank details must be near-perfect, descriptions less so); latency (batch is fine: hours, not seconds); languages; how much human review capacity exists.
- Pipeline: ingest → classify (invoice vs credit note vs junk) → text and layout extraction (native PDF text where available, OCR or a vision-language model for scans) → LLM extraction into a JSON schema with structured outputs (constrained decoding guarantees schema-valid JSON, but not correct values Anthropic docs) → deterministic validation (line items sum to the subtotal, subtotal plus tax equals total, dates parse, supplier and PO number match the vendor master and open POs) → confidence-based routing to human review → ERP.
- Numbers: 2M/month ≈ 67k/day ≈ 0.8/s, so a batch workload. At roughly 3k input tokens (page text or image) and 500 output tokens per invoice on a small model at batch price (say $0.50 / $2.50 per M), it's about $0.003 per invoice, or ~$6k/month. Images cost more tokens than text, so use extracted text when it's reliable. Human review is the real cost: if 10% go to review at 2 minutes each, that's 13k minutes/day, about 28 full-time reviewers.
- Key insight: the business metric is "straight-through processing rate at a target error rate." Accounting checks catch most extraction errors for free; use them as both a validator and a confidence signal. Per-supplier templates or few-shot examples (retrieved by supplier ID) lift accuracy for high-volume suppliers.
- Evals: field-level exact match on a labeled set stratified by supplier and document quality; track error rate per field among auto-approved invoices (the dangerous errors); reviewer corrections become new labels.
- Deep dives: prompt injection embedded in documents ("pay to account X") is a real fraud vector, so bank details must be validated against the vendor master and changes must go through a human; multi-page line items; when to fine-tune a small extractor (once you have hundreds of thousands of reviewer-corrected labels).
Sketch C: Content moderation with LLMs
Prompt: "Design a moderation system for user posts on a large social platform, using LLMs."
- Clarify: volume (assume 50M posts/day ≈ 580/s average); modalities (text, images, video); the policy taxonomy and how often it changes; latency (pre-publish blocking vs post-publish takedown); appeals; regional legal requirements; the cost of false positives (silencing users) vs false negatives (harm).
- Cascade architecture, because running a frontier LLM on every post is too expensive: (1) hash matching for known illegal content and spam fingerprints; (2) cheap, fast classifiers (small fine-tuned models) on everything, handling the clear majority; (3) an LLM policy reasoner on the uncertain middle band, given the written policy and the content, returning a label plus rationale; (4) human reviewers for the hardest cases, appeals and high-severity categories.
- LLM role: safety classifiers built on LLMs, like Llama Guard, classify both prompts and responses against a risk taxonomy and can be adapted to new taxonomies via instructions. Inan+ 2023 More recent "policy as prompt" reasoning models, such as OpenAI's open-weight gpt-oss-safeguard (October 2025), read a developer-written policy at inference time and explain their decision, so policy changes don't require retraining. OpenAI 2025 LLMs are also useful offline: labeling training data for the cheap classifiers, and drafting policy clarifications from reviewer disagreements.
- Numbers: if 5% of 50M posts reach the LLM tier at ~800 tokens each, that's 2B tokens/day; at a small model price (~$1/M input) that's about $2k/day, versus ~$40k/day to send everything. The cascade is the cost design.
- Evals: precision and recall per policy category on a golden set labeled by policy experts; inter-annotator agreement as the ceiling; drift monitoring (new slang, coordinated evasion); fairness checks across dialects and languages, where classifiers often have higher false-positive rates.
- Deep dives: adversarial evasion (leetspeak, images of text, coded language); threshold tuning per category and per region; reviewer wellbeing (blur, limit exposure); transparency and appeal flows.
All three sketches reward the same move: put cheap deterministic checks and small models in front, reserve the expensive LLM for the uncertain middle, and keep humans for the highest-stakes decisions. If you can draw that cascade and do the "what fraction reaches each tier" math, you've answered most of the question.
- RouteLLM (Ong et al., 2024): learned routing between strong and weak models.
- LiteLLM: Fallbacks: concrete gateway failover patterns.
- Llama Guard (Inan et al., 2023): an LLM-based input/output safety classifier.
- OpenAI: Introducing gpt-oss-safeguard: policy-as-prompt safety reasoning models.
Cross-cutting patterns: a cheat sheet
The six case studies reuse a small set of patterns. If you can name these and say when each applies, you can improvise a design for an unfamiliar prompt.
| Pattern | Used in | When to reach for it |
|---|---|---|
| Router + fixed workflows for the head, agent for the tail | Support, moderation, gateway | Traffic is skewed toward a few repetitive intents. |
| Deterministic tool layer (auth, limits, idempotency) | Support, voice, coding | Any side effect. Always. |
| Hybrid retrieval + rerank + permission pre-filter | Enterprise search, memory, support policy | Knowledge that changes or is access-controlled. |
| Orchestrator-workers with context isolation | Deep research, coding (sub-agents for exploration) | Breadth-first, separable subtasks that each read a lot. |
| Verifier in the loop (tests, accounting checks, citation checks) | Coding, extraction, research | A cheap automatic check exists. Build one if it doesn't. |
| Cascade (cheap → expensive → human) | Moderation, extraction, support escalation | High volume, most items easy, a few hard and high-stakes. |
| Async extraction + consolidation | Memory, ingestion | Work that doesn't have to be on the latency path; use batch pricing. |
| Streaming everywhere + speculative execution | Voice, chat UIs | Perceived latency matters more than total time. |
| Durable execution with checkpoints and budgets | Coding, research | Long-running agents that must survive failures and not run away. |
| Break the lethal trifecta | Every agent that reads untrusted content | Never combine private data, untrusted input and an exfiltration channel without a hard control on at least one. Willison 2025 |
| Case study | Binding constraint | Dominant cost | Primary eval signal | Scaling unit |
|---|---|---|---|---|
| Support agent | Correctness of actions, policy compliance | LLM input tokens (agent loop) | Database end state + resolution rate | Conversations/day, model QPS |
| Enterprise RAG | Permissions, retrieval quality | Generation tokens | Recall@k + citation faithfulness + leak tests | Corpus size, query QPS |
| Coding agent | Verifiability, sandbox security | Long-loop input tokens (caching critical); reviewer time | Tests pass + merge rate | Concurrent sandboxes |
| Deep research | Source and citation quality | Tokens across many subagents | Rubric judge + citation checks | Concurrent agent loops, provider TPM |
| Voice agent | Latency (voice-to-voice) and turn-taking | TTS + telephony, often more than the LLM | Task completion + latency percentiles | Concurrent calls |
| Memory assistant | Privacy, correctness over time | Injected memory tokens on every call | Extraction precision, update correctness, appropriateness | Users × memories (partitioned) |
Interview question bank
1. You're asked to "design a chatbot for our docs." What are your first five questions, and why?
(1) Who are the users and what are they trying to do: developers debugging, or prospects evaluating? That sets tone, depth and the success metric. (2) What does a correct answer look like, and how do we know: is there a support-ticket history to build a golden set from? (3) How big and how fresh is the corpus, and are there access restrictions (public docs vs customer-specific content)? (4) What are the latency and cost constraints: streaming chat with under 2 s to first token, and a cost ceiling per query or per month? (5) What happens when the bot can't answer: link to search, open a ticket, hand off to a human? Each question changes the design. Freshness and access drive the ingestion design, the success definition drives the eval, and the fallback drives the escalation path. Close by proposing a baseline (hybrid retrieval plus one grounded prompt with citations) and the metric you'd use to decide whether to go further.
2. How is AI system design different from classic system design? Give three concrete differences.
First, correctness is statistical: the same input can yield different outputs and a successful HTTP response can still be wrong, so the eval set becomes the spec and every change is judged by running it. Second, unit cost is high and variable: cents to dollars per request, scaling with tokens, so cost is a primary design axis and you do token math like you'd do QPS math. Third, the main trade-off is quality versus latency versus cost: bigger models, more context, more agent steps and more samples all buy quality at the expense of the other two. Additional differences worth naming: a new threat model (the component reads attacker-controlled text and might follow it), dependence on external model providers with rate limits and silent version changes, and latency dominated by token generation rather than I/O.
3. Back-of-envelope: a support bot handles 50k conversations/day, 12 model calls each, ~8.5k input and 250 output tokens per call, at $3/$15 per M. What's the daily cost, and how would you halve it?
Per conversation: 12 × 8.5k = 102k input tokens ≈ $0.31, and 12 × 250 = 3k output ≈ $0.045, so about $0.35. Times 50k ≈ $17.5k/day. To halve it: (1) prompt-cache the static ~5k-token prefix (system prompt, policies, tool schemas). At about 0.1× input price for cache reads, that alone brings it to roughly $0.20/conversation, around $10k/day. (2) Route the easy 40% of intents (order tracking, FAQ) to a small model or a fixed workflow at a third of the price, saving roughly another 25%. (3) Trim tool outputs to the fields the model needs, and cap retrieved context. Then verify with the eval suite that resolution rate and policy compliance didn't drop. Cost cuts that hurt quality aren't savings.
4. When should you use an agent instead of a fixed workflow?
Use a fixed workflow when the steps are known in advance and the variation is in the content, not the procedure: classify, retrieve, answer; or extract, validate, route. Workflows are cheaper, faster, more predictable and easier to evaluate. Use an agent when the number and order of steps depend on what's discovered along the way and can't be enumerated: debugging code, open-ended research, multi-step support issues with unpredictable tool needs. The cost is higher latency, more tokens, more variance and harder evals, so the bar is an eval showing the workflow fails for reasons the agent fixes. In practice, the best designs are hybrid: a router sends the predictable head of traffic to workflows and the long tail to an agent, and even the agent gets a suggested plan structure.
5. How do you prevent a support agent with refund tools from being socially engineered or prompt-injected into issuing bad refunds?
Defense in depth, with the critical controls outside the model. Identity is bound server-side from the authenticated session, so the model can't choose whose orders to act on. Tools are narrow and validate everything: ownership, eligibility windows, amount limits, fraud scores, via a deterministic policy engine. Mutations above a threshold go to a human approval queue; all mutations use idempotency keys and are logged. The model is told to treat tool outputs and customer text as data, but you assume that will sometimes fail, which is why the tool layer enforces policy regardless of what the model requests. A confirmation step before mutations catches honest confusion. Finally, an adversarial eval set (fake manager approvals, injected instructions in order notes) runs in CI, and production monitoring alerts on refund-rate anomalies.
6. In enterprise RAG, where do you enforce document permissions, and why not just post-filter results?
Enforce them inside the retrieval query: each chunk is indexed with the principals (users and groups) allowed to see it, and the user's principal set is a filter in the vector and keyword search. Then re-check the final handful of documents against the source system's API, cached briefly, to cover sync lag after revocations. Post-filtering is worse for three reasons: if most of the top-k is forbidden you return few or no results (or must over-fetch massively), it can leak information through counts or timing, and it's easy to get wrong in one code path. Never rely on the prompt to enforce permissions: once content is in the context window it can appear in the answer. Mention the revocation SLA and how you'd test it (a red-team suite with zero-leak pass criteria).
7. Users complain the docs assistant gives wrong answers. How do you debug whether it's retrieval or generation?
Pull the traces for the failing queries and check whether the chunk containing the right answer was in the retrieved context. If it wasn't, it's a retrieval problem: look at chunking (was the answer split across chunks?), query formulation (conversational follow-ups not rewritten), lexical vs semantic mismatch (exact identifiers need BM25), missing or stale documents, or permission filters excluding it. If it was retrieved but the answer was still wrong, it's generation: the model ignored it, it was buried mid-context, conflicting documents confused it, or the prompt didn't require grounding. Measure the two separately going forward: recall@k on a labeled retrieval set, and faithfulness and correctness given gold context. Most RAG issues turn out to be retrieval issues, so fix those first.
8. Estimate the vector storage for 40M chunks with 1024-dim embeddings. What are your options to shrink it?
Float32: 40M × 1024 × 4 bytes ≈ 164 GB, plus index overhead (HNSW graphs can add a large fraction on top). Options: int8 scalar quantization (~41 GB) with small recall loss; binary quantization (1 bit per dimension, ~5 GB) used for a fast first pass with rescoring on full-precision vectors; product quantization; smaller embedding dimensions (some models are trained so that truncated "Matryoshka" prefixes still work, e.g. 256 dims); and deduplicating near-duplicate chunks, which is common in enterprise corpora. The point to make is that storage is rarely the cost bottleneck at this scale; generation tokens and retrieval quality matter more. It's a sizing question, not a blocker.
9. Why does prompt caching matter so much for coding agents specifically?
Agent loops re-send the whole conversation each step: the system prompt, tool definitions, every file read and every command output so far. If a task runs 50 steps with context growing to 80k tokens, the total input is around 2M tokens, and without caching the cost grows roughly quadratically with step count. Since each step's prompt is the previous prompt plus a small addition, almost all of it is a cacheable prefix. At cache-read prices around a tenth of normal input, the per-task cost drops by something like 3–4× in my earlier estimate (about $6 to about $1.60), and latency to first token also improves. Practical implications: keep the prefix stable (don't put timestamps or random IDs at the top), append rather than rewrite history, and be aware that compaction invalidates the cache, so compact rarely and deliberately.
10. How would you evaluate an autonomous coding agent before rolling it out internally?
Build an internal benchmark from your own history: merged PRs that fixed an issue and included tests. Reset each repo to the parent commit, give the agent the issue text, and grade by running the PR's tests (tests that should go from failing to passing, plus the existing suite to catch regressions), all inside the same sandbox the agent will use in production. Report resolution rate, cost and time per task, and process metrics (did it run tests before finishing, how many steps). Add a human review of a sample for mergeability: minimal diff, idiomatic, tests added, no test gaming. Use public SWE-bench variants only as a sanity check, since they may be contaminated or unrepresentative of your stack. Then roll out to volunteer teams with draft PRs, and track merge rate, reviewer time and post-merge reverts.
11. What can go wrong if a coding agent's sandbox has open internet access and the repo's secrets?
That combination is the lethal trifecta: private data (secrets, proprietary code), exposure to untrusted content (issue text, code comments, dependency READMEs or web pages the agent reads), and an exfiltration channel (outbound network). A prompt injection planted in any of that content could instruct the agent to send secrets to an attacker-controlled URL, and the agent may comply because models can't reliably distinguish instructions from data. Fixes: no production secrets in the sandbox at all; short-lived, narrowly scoped tokens (push only to agent branches in one repo); network egress allowlisted to package mirrors; microVM or gVisor isolation; and human merge as the final gate. Remove any one leg of the trifecta with a hard control, ideally more than one.
12. Why did Anthropic's research system use multiple agents, and when would you argue against multi-agent?
Research is breadth-first and parallelizable: subtopics can be investigated independently, and quality scales with how much material is read. Separate subagents each get a fresh context window, read lots of pages, and return compressed findings; they also run in parallel, which cut research time by up to 90% on complex queries. Anthropic reported the multi-agent version outperformed a single agent by 90.2% on its internal eval, and that token usage explained 80% of performance variance on BrowseComp. Against: it uses roughly 15× the tokens of a chat; subagents can duplicate work or make conflicting assumptions; it's harder to debug. For tightly coupled tasks like most coding, where every decision depends on shared context, a single-threaded agent (possibly with read-only exploration sub-agents) is usually better, which is the argument Cognition made.
13. Back-of-envelope: what does a deep-research report cost, and what drives it?
Assume a lead agent plus 5 subagents, each subagent making about 20 tool calls with ~30k tokens of context per step: about 3M input tokens, plus ~200k for the lead and ~100k for a citation pass, for ~3.3M input and ~70k output. At $3/$15 per M that's about $10 + $1 ≈ $11 uncached, and caching within each subagent loop brings it to roughly $4–6, plus ~$0.75 for 75 searches at $10/1,000. The drivers are the number of subagents, steps per subagent, and the size of fetched pages kept in context. Levers: scale effort to query complexity (one agent for simple questions), extract only relevant passages from fetched pages, use cheaper models for workers, and cache. At 5,000 reports/day that's tens of thousands of dollars daily, which is why these features are often rate-limited.
14. Walk me through the latency budget of a cascaded voice agent. Where would you cut first?
A typical breakdown (from a widely used practitioner primer): mic and client buffering ~40 ms, network and codecs ~140 ms round trip, transcription plus endpointing ~300 ms, LLM time to first token ~650 ms, TTS time to first byte ~120 ms, totaling around 1.3 s. Phone calls add telephony buffering on top. The biggest items are LLM TTFT and endpointing, so cut there first: a faster model with a short, cached prompt and no reasoning mode on the hot path; a turn-detection model instead of a long silence timeout; streaming at every stage (TTS starts on the first sentence); co-locating STT, LLM and TTS in one region; and speculative generation on confident partial transcripts. For slow tool calls, play a short filler phrase so the caller isn't met with silence.
15. The user interrupts the voice agent mid-sentence. What exactly should happen?
VAD detects user speech while the agent is speaking. After a short confirmation (a minimum duration or a few transcribed words, so a cough or "mm-hm" doesn't trigger it), the system stops TTS playback immediately, flushes audio already buffered downstream (for example, a "clear" message on Twilio Media Streams), and cancels the in-flight LLM generation and any queued TTS. Then the context is fixed: the assistant message in history is truncated to what the caller actually heard, using TTS word timestamps or playback marks, so the model doesn't assume the user received information they didn't. Then the system listens to the new utterance and responds to it. Measure false barge-ins and missed barge-ins as explicit metrics.
16. Cascaded STT→LLM→TTS vs speech-to-speech for a clinic booking line: which, and why?
Cascaded, for now. The booking line is tool-heavy (search slots, book, cancel), compliance-heavy (transcripts for audit, identity verification before disclosing health information), and cost-sensitive at high volume. The cascaded pipeline gives a full text transcript for every turn, lets you pick the best STT for medical vocabulary and the most reliable tool-calling text model, makes context management and evals straightforward, and has been estimated at a fraction of speech-to-speech cost. Speech-to-speech wins on naturalness and can win on latency, and it hears tone, so it's attractive for companionship, coaching or language practice. I'd revisit the decision periodically, because speech-to-speech models have been improving quickly, and a hybrid is plausible.
17. How do you handle a memory system when user facts change ("I'm vegetarian" → "I eat fish now")?
Don't append blindly. The extraction step produces a candidate fact, and a consolidation step retrieves the user's semantically nearest existing memories and decides add, update, delete or no-op. Here it's an update: the old memory is marked superseded (with a valid-to timestamp and a pointer to the new one), not physically overwritten, so the assistant can reason about time and you can roll back bad updates. Prefer explicit user statements over inferences, record confidence, and show a "memory updated" indicator so the user can correct mistakes. Evaluate with scripted multi-session scenarios where facts change, checking the latest fact wins, which is the "knowledge updates" ability in LongMemEval.
18. What are the privacy risks of a long-term memory assistant, and how do you design for them?
Risks: storing sensitive categories (health, sexuality, religion) without meaningful consent; surfacing memories in contexts where they feel intrusive; cross-user leakage; memory poisoning, where injected content in a web page or document gets "remembered"; incomplete deletion under GDPR or CCPA; and memories about third parties who never consented. Design: category whitelists with opt-in for sensitive ones; extraction only from the user's own messages, not tool outputs; user ID enforced in the storage layer with per-user encryption; a visible, editable memory list with a temporary-chat mode and an off switch; sensitivity-aware retrieval rules; and a deletion pipeline that tracks lineage to embeddings, summaries, caches and backups, with automated tests that verify deletion end to end.
19. Design question: build a learned router that sends queries to a cheap or expensive model. How do you train and validate it?
Start with data: log production queries, run a sample through both models, and label which response is acceptable (with a calibrated LLM judge, spot-checked by humans) or which is preferred. Train a lightweight classifier (an embedding plus a small head, or a small fine-tuned model) to predict "the cheap model is good enough" for a query, as in RouteLLM's preference-data approach. Choose the threshold on a held-out set to hit a target quality, such as no more than a 1-point drop on your eval, and read the cost saving off the resulting routing fraction. Validate with an online A/B on the business metric, and keep a small random sample going to the expensive model so you can detect drift. Watch for distribution shift (new query types the router has never seen) and per-segment regressions hidden by averages.
20. Invoice extraction: how do you decide which invoices skip human review?
Combine independent signals into a routing rule, and tune it against a labeled set for a target error rate among auto-approved items. Signals: deterministic validation (line items sum to subtotal, subtotal plus tax equals total, dates valid, currency consistent), cross-checks against business data (supplier exists in the vendor master, PO number matches an open PO with matching amounts, bank details unchanged), model confidence (field-level log-probabilities, or agreement between two extraction passes or models), and supplier history (error rate on this supplier's past invoices). Anything failing a hard check, involving changed bank details, or above an amount threshold goes to review regardless of confidence. Track the straight-through rate and the error rate among auto-approved invoices per field, and feed reviewer corrections back as labels.
21. Your LLM provider silently updates the model and quality drops. How would you detect and prevent this?
Prevent: pin model versions explicitly rather than using floating aliases, and treat model upgrades like dependency upgrades, gated by the regression eval suite in CI. Detect: run a canary eval set continuously against production configurations (daily or hourly), and monitor online proxies such as thumbs-down rate, escalation rate, tool-error rate, output length distributions and refusal rates, with alerts on shifts. Have the gateway log the exact model version per request, so you can correlate. Mitigate: a fallback to the previous version or another provider via the gateway, and a prompt or eval update process when you deliberately migrate. Provider deprecation schedules mean you'll be forced to migrate eventually, so a fast eval-driven migration path is part of the design.
22. How would you size infrastructure for 1,000 concurrent voice calls?
Concurrency, not QPS, is the unit. Each call holds a stateful session: a WebSocket or WebRTC connection, audio buffers, VAD and turn-detection inference (small, CPU-friendly models at tens of milliseconds per inference), and streaming connections to STT and TTS. Benchmark how many sessions one worker instance handles at target latency (CPU-bound on audio processing and model inference), then provision for peak plus headroom and spread across regions near callers and providers. Check provider concurrency limits for STT, TTS and the LLM (tokens per minute at roughly 12k input tokens per call-minute in my earlier estimate), and negotiate them in advance. Autoscale on active sessions, drain gracefully on deploys (calls can't be migrated mid-stream easily), and keep a fallback path to human agents or voicemail when capacity is exhausted.
23. What does an iteration plan look like after launching any of these systems?
Phase 1: shadow mode or an internal dogfood period, where the system runs but humans act, to compare decisions. Phase 2: a canary to a small percentage of traffic with tight monitoring and an instant rollback switch. Then ramp with an A/B against the baseline on the primary metric and guardrail metrics. Throughout, log full traces; sample failures (thumbs-down, escalations, judge-flagged answers) weekly; triage them into categories; turn representative failures into new eval cases; and fix with the cheapest lever first (prompt or tool description, then retrieval, then model or architecture changes). Re-run the full eval before every change ships. Later, once you have volume and labeled successes, consider distillation or fine-tuning a smaller model for the highest-volume paths to cut cost and latency.