Agent Architectures
An "agent" is an LLM running in a loop, choosing tools and acting on what comes back, until it decides it's done or a limit stops it. Almost every AI engineering interview now has an agent question, either conceptual ("ReAct vs plan-and-execute?"), design-shaped ("design a research agent / support agent / coding agent"), or skeptical ("why not just a workflow?"). This page covers the whole design space: the workflow patterns that should come first, the single-agent loop and its reasoning variants, planning, the main multi-agent topologies and the 2025–2026 argument over when they pay off, plus state, durability, protocols, frameworks and failure modes. The goal is that you can pick an architecture for a given problem and defend it with mechanisms and numbers.
TL;DR: the 8–12 things to be able to say out loud
- Agent = LLM + tools + loop + state. The model picks the next action, the harness runs it, and the result goes back into context. Autonomy is a spectrum, from a single call to a fixed workflow to a router to a full agent loop.
- Workflows first. Anthropic's taxonomy: prompt chaining, routing, parallelization (sectioning / voting), orchestrator–workers, evaluator–optimizer. Use an agent only when you can't fix the steps in advance and the extra cost and variance are worth it.
- The loop needs brakes: max steps, token/$ budgets, wall-clock timeouts, no-progress detection and an explicit "finish" tool. Unbounded loops are the classic production incident.
- ReAct interleaves thought → action → observation. Plan-and-execute / ReWOO plan up front, which saves tokens and latency but is brittle. LLMCompiler plans a DAG and runs it in parallel. Reflexion / evaluator loops add self-critique. ToT / LATS add search. CodeAct uses code as the action space.
- Planning helps on long-horizon, decomposable tasks, as long as the plan is a living artifact (todo list / DAG / file) that gets revised as observations come in.
- Multi-agent topologies: orchestrator–workers, supervisor/router, hierarchical, handoffs/swarm, pipeline, debate, blackboard, market/peer-to-peer. Each trades context isolation and parallelism against coordination cost and lost context.
- The debate: Cognition (2025) said "don't build multi-agents" because actions carry implicit decisions and subagents lack shared context. Anthropic (2025) reported big gains from parallel subagents on breadth-first research at roughly 15× the tokens of chat. Synthesis: parallelize reading, keep writing single-threaded. Cognition's own 2026 follow-up lands in the same place.
- Handoff context is the hard part: pass a structured task spec (objective, constraints, output format, budget, what's already known), not a vague one-liner.
- State and control flow: graph state machines (LangGraph-style), checkpoints per step, human-in-the-loop interrupts, resumability. For long-running work, use durable execution (Temporal-style), idempotent tools and sagas.
- Protocols: MCP = agent ↔ tools/context. A2A = agent ↔ agent across vendors. Both now sit under Linux Foundation governance.
- Failure modes: loops, tool errors, compounding error (\(p^n\)), goal drift, context overflow, over-delegation, cost blowups. The mitigations are mostly ordinary engineering: budgets, validation, verification, observability, evals.
1. What an agent is
Here's the working definition most practitioners use: an agent is a system where an LLM dynamically directs its own process and tool usage, keeping control over how it accomplishes the task. A workflow is one where LLMs and tools are orchestrated through predefined code paths. Anthropic 2024 Both are "agentic systems". What separates them is who decides the control flow: your code or the model.
Mechanically, an agent has four parts:
LLM (the policy)
Given the current context (instructions, history, observations), it emits either a final answer or one or more tool calls (structured JSON naming a tool plus arguments). It's a stochastic policy \( \pi_\theta(a_t \mid c_t) \) over actions.
Tools (the action space)
Functions the harness can execute: search, code execution, file edits, APIs, DB queries, other agents. Tool descriptions are part of the prompt, and their quality largely determines how reliable the agent is.
Loop (the harness)
Code that calls the model, executes the requested tools, appends results, and repeats until a stop condition. The model never "runs" anything itself. The harness does, so the harness is where you put permissions, budgets and retries.
State (memory)
The context window (short-term), plus any external state: scratchpads, todo files, vector stores, checkpoints. See Memory. State is what lets the agent resume, hand off, or survive context limits.
The basic building block is the augmented LLM: one model call that can use retrieval, tools and memory. Anthropic 2024 Every pattern below is a composition of augmented-LLM calls with different control flow around them.
The autonomy spectrum
"Agent or not" is the wrong binary. Think of a dial for how much control flow the model owns:
| Level | Who decides control flow | Example | Predictability | Typical cost / latency |
|---|---|---|---|---|
| 0. Single call | Nobody (one shot) | Summarize this doc; extract fields | Highest | 1 call |
| 1. Chain / workflow | Your code, fixed steps | Outline → draft → translate | High | k calls, fixed |
| 2. Router | LLM picks a branch; code defines branches | Classify ticket → billing / tech / refund flow | High | 1 + branch |
| 3. Tool-calling loop | LLM picks tools and when to stop, within one task | Support agent with lookup/refund tools | Medium | variable, often 3–20 calls |
| 4. Planner + executor / multi-agent | LLM decomposes the task, spawns sub-work | Deep research, large refactors | Low | tens to hundreds of calls |
| 5. Long-running autonomous | LLM sets sub-goals across sessions; humans check in | Background coding agents working for hours | Lowest | hours, many context windows |
Each step up the dial buys flexibility on open-ended inputs and pays for it in variance, cost, latency and debuggability. The engineering question is always: what's the lowest autonomy level that handles the input distribution I actually have? Most production "agents" are levels 1–3 wearing a level-4 marketing label.
"What is an agent?" is a warm-up, and they're checking whether you separate the model from the harness. A strong answer: an LLM in a loop that emits tool calls; a harness that executes them, enforces permissions and budgets, and feeds back observations; state that persists across steps; and a stop condition. Add the workflow vs agent distinction (who owns control flow) and say you'd default to the simplest level that works.
- Anthropic: Building effective agents (Dec 2024): the canonical workflow/agent taxonomy; short and practical.
- Lilian Weng: LLM Powered Autonomous Agents (2023): the planning / memory / tools decomposition that most courses still use.
- OpenAI: A practical guide to building agents (PDF): OpenAI's view on when to build agents, guardrails and orchestration.
2. Workflows vs agents: the five workflow patterns
Anthropic's "Building effective agents" post (Dec 2024) became the shared vocabulary for this. It lists five workflow patterns and then the autonomous agent as the sixth, most flexible option. Anthropic 2024 Interviewers expect you to know the names and when each one fits.
The patterns, mechanically
Prompt chaining. Break the task into fixed sequential steps; each LLM call consumes the previous output. Add programmatic gates between steps (schema validation, length check, a classifier) so a bad intermediate result fails fast instead of contaminating later steps. You spend latency to get accuracy: each call does a simpler job. Use it when the decomposition is known and stable (generate marketing copy → translate it; write an outline → check it against criteria → write the doc).
Routing. Classify the input, then dispatch to a specialized prompt, toolset or model. This gives separation of concerns: optimizing the refund prompt can't degrade the tech-support prompt. It also handles cost routing (easy queries to a small model, hard ones to a frontier model). The router can be an LLM, a fine-tuned classifier or an embedding nearest-neighbor. Routing errors are silent, so log the route and evaluate the router on its own.
Parallelization. There are two flavors. Sectioning splits a task into independent subtasks that run concurrently (one call answers the user while another screens for policy violations; one call per document section). Voting runs the same task several times and aggregates. That's self-consistency applied at the system level Wang+ 2022, useful for reviews ("flag if any of 3 reviewers flags") or high-stakes classification. With sectioning, latency is the slowest branch instead of the sum. With voting, cost scales linearly with the number of votes.
Orchestrator–workers. A central LLM decomposes the task at runtime, delegates subtasks to workers, and synthesizes the results. The difference from parallelization is that the subtasks aren't predefined; they depend on the input (which files need changing, which sub-questions need researching). This is already a small multi-agent system, and §6 develops it further.
Evaluator–optimizer. One call generates, another evaluates against explicit criteria and returns feedback, and the loop repeats until it passes or hits a max-iteration cap. It works when (a) clear evaluation criteria exist and (b) feedback measurably improves the output, the way a human editor's would. Literary translation, code that must pass tests, and search tasks that need several rounds are good fits. It's the workflow version of Self-Refine Madaan+ 2023 and Reflexion (§4).
| Pattern | Use when | Avoid when | Cost / latency shape |
|---|---|---|---|
| Prompt chaining | Fixed, known decomposition; each step checkable | Steps depend heavily on the content | Sum of k calls; sequential latency |
| Routing | Distinct input categories needing different handling; cost tiers | Categories overlap heavily; misroutes are expensive and invisible | +1 cheap call; can reduce total cost |
| Parallel: sectioning | Independent subtasks; guardrails alongside main task | Subtasks share hidden dependencies | Cost = sum; latency = max |
| Parallel: voting | Need confidence or recall; cheap-ish calls | Correlated errors (all samples wrong the same way) | n× cost; latency ≈ 1 call |
| Orchestrator–workers | Subtasks unknown until runtime (multi-file edits, research) | The decomposition is actually fixed (use chaining) | Variable; 1 plan + n workers + synth |
| Evaluator–optimizer | Clear criteria, iterative improvement measurable | Evaluator no better than generator (self-grading echo chamber) | 2× per iteration × iterations |
| Agent | Open-ended, unpredictable number of steps, trusted environment, feedback available | Low tolerance for variance; steps knowable | Unbounded unless budgeted |
Reaching for a multi-agent framework when a 3-step chain with validation gates would do. Interviewers notice. The canonical advice is to start with direct API calls and the simplest composition, and add complexity only when it demonstrably improves outcomes on your evals. Frameworks "often create extra layers of abstraction that can obscure the underlying prompts and responses". Anthropic 2024
A typical probe: "Design a system to answer customer emails." The weak answer goes straight to "a multi-agent system with a planner…". A strong answer starts with routing (intent classification → per-intent chain), adds parallel sectioning for a policy/PII guardrail, uses evaluator–optimizer for tone/accuracy on drafts, and keeps a narrow tool-calling agent only for the long-tail "investigate account" intent. Then it says how you'd measure whether the agent branch earns its cost.
- Anthropic: Building effective agents: the source taxonomy plus an appendix on tool (ACI) design.
- LangChain docs: multi-agent patterns: subagents vs handoffs vs skills vs router, with call/token counts per pattern.
- Self-Refine (Madaan+ 2023): the generate–feedback–refine loop, measured.
3. Anatomy of the agent loop
Every agent framework, however much it abstracts, implements something like the loop below. You should be able to write it on a whiteboard and point out where each production concern lives.
def run_agent(task, tools, llm, limits):
messages = [system_prompt(tools), user(task)]
state = State(steps=0, tokens=0, cost=0.0, started=now(), seen_actions=Counter())
while True:
# ---- budget / stop checks (harness-enforced, NOT model-enforced) ----
if state.steps >= limits.max_steps: return finalize(messages, reason="max_steps")
if state.cost >= limits.max_cost_usd: return finalize(messages, reason="budget")
if now() - state.started > limits.timeout: return finalize(messages, reason="timeout")
if context_tokens(messages) > limits.ctx_soft_cap:
messages = compact(messages) # summarize / drop old tool outputs
# ---- THINK: model proposes next action(s) ----
resp = llm(messages, tools=tools) # may contain text + 0..n tool calls
state.update_usage(resp.usage)
messages.append(assistant(resp))
if not resp.tool_calls: # model chose to answer
return resp.text # (or require an explicit finish() tool)
# ---- ACT: execute tool calls (parallel if independent) ----
results = []
for call in resp.tool_calls: # or asyncio.gather for parallel calls
if not policy.allows(call): # permissions / HITL approval
results.append(tool_error(call, "denied by policy")); continue
state.seen_actions[hash(call)] += 1
if state.seen_actions[hash(call)] > 2: # loop detection
results.append(tool_error(call, "repeated identical call; try something else")); continue
try:
out = tools[call.name].run(validate(call.args), timeout=limits.tool_timeout)
results.append(tool_result(call, truncate(out, limits.max_obs_tokens)))
except Exception as e: # errors go BACK to the model as observations
results.append(tool_error(call, str(e)))
# ---- OBSERVE: feed results back; loop ----
messages.extend(results)
state.steps += 1
checkpoint(state, messages) # durability / resumability
Stop conditions and budgets
An agent ends for one of two kinds of reasons: it says it's done (natural), or the harness stops it (forced). You need both.
| Stop condition | Mechanism | Notes |
|---|---|---|
| Natural finish | Model returns text with no tool call, or calls an explicit finish(answer) / submit tool | An explicit finish tool with a schema makes completion detectable and the output structured. Models sometimes "finish" prematurely, so pair it with verification. |
| Max steps / turns | Counter in harness | Typical caps range from about 10 (support) to hundreds (coding). Pick from the eval-trace distribution, e.g. p99 steps of successful runs × 1.5. |
| Token / $ budget | Sum usage per call | Cost grows superlinearly with steps because each call resends the history (see below). Prompt caching helps a lot. |
| Wall-clock timeout | Deadline | Essential for user-facing latency SLOs and for detecting hung tools. |
| No-progress / loop detection | Hash recent (tool, args); detect repeated errors; compare state snapshots | Respond by injecting a nudge ("you've tried X 3 times"), escalating, or aborting. |
| Verification gate | Tests pass, schema valid, judge approves | Turns "I think I'm done" into "the environment confirms done". |
| Human interrupt | Approval required for risky actions; user cancels | See §10 on HITL. |
The quadratic cost of loops
Each step resends the whole history. If every step adds about \(\Delta\) tokens (thought + tool call + observation) on top of a base prompt \(P\), then step \(t\) has input \(P + t\Delta\), and total input tokens over \(n\) steps are
$$ \text{Input}_{\text{total}} \approx nP + \Delta \frac{n(n-1)}{2} = O(nP + n^2\Delta). $$Back-of-envelope: \(P = 5\text{k}\), \(\Delta = 2\text{k}\), \(n = 30\) gives \(150\text{k} + 870\text{k} \approx 1\text{M}\) input tokens for one task. That's why (a) prompt caching matters so much for agents, since the stable prefix is re-read at a large discount; (b) truncating or summarizing tool observations pays off (\(\Delta\) is usually dominated by tool output); and (c) subagents can be cheaper per useful token than one long thread, because their large intermediate observations never enter the parent's context.
Throwing on tool errors. A failed tool call should usually go back to the model as an observation (with an actionable message: "file not found: use an absolute path; did you mean /src/app.py?") so it can self-correct. Only harness-level faults (budget, policy, crash) should end the loop. Equally bad: returning 50k tokens of raw HTML or logs as an observation. Truncate, paginate, or summarize.
"Your agent sometimes runs forever / costs $40 on one request. What do you do?" Strong answers stack several layers: hard caps (steps, $, time) in the harness; loop detection on repeated identical calls; observation truncation and context compaction; prompt caching; an explicit finish tool plus verification; per-tenant budgets and alerts; and evals that track the step-count distribution so regressions show up before production.
- Claude Agent SDK overview: a production harness (Claude Code's loop) exposed as a library: tools, hooks, permissions, sessions.
- Anthropic: Writing effective tools for agents: tool design, response truncation, actionable error messages.
- Anthropic: Effective context engineering for AI agents: compaction, note-taking and subagents as context-management tools.
4. Single-agent reasoning patterns
These patterns differ in when the model reasons relative to acting, how much it plans ahead, and whether it explores alternatives. Most came from 2022–2024 papers. Today's models have absorbed many of them natively (reasoning models "think" between tool calls; frontier models emit parallel tool calls), so treat them as design vocabulary rather than recipes you have to implement by hand.
ReAct: interleaved reasoning and acting
ReAct prompts the model to alternate Thought → Action → Observation steps, so reasoning traces guide which tool to call next and observations ground the following reasoning. Yao+ 2022 The paper showed that combining the two beat reasoning-only chain-of-thought (which hallucinates facts) and acting-only (which can't plan or recover) on QA/fact-checking and interactive tasks. ReAct is the default shape of essentially every tool-calling loop today. Native function calling replaced the text "Action:" parsing, and reasoning models put the "Thought" in hidden thinking tokens.
- Strengths: adaptive, since each step uses the latest observation; robust to surprises; simple.
- Weaknesses: one LLM call per action (latency, cost); greedy, with no global plan, so it can wander or loop; history grows quadratically in cost.
Plan-and-execute / Plan-and-Solve
First generate a full plan (a list of steps), then execute the steps, often with a cheaper executor model or plain code, and optionally replan after each step or on failure. Plan-and-Solve prompting showed that asking the model to "first devise a plan, then carry it out" reduces missing-step errors compared with zero-shot CoT. Wang+ 2023 In agent systems the pattern is: a planner (strong model) emits steps; an executor (ReAct sub-loop, or a smaller model) runs each one; a replanner looks at results and revises the remaining steps.
- Strengths: global view of the task; a big model is called rarely; the plan is inspectable (and human-approvable) before execution.
- Weaknesses: plans go stale when early observations contradict assumptions, so you need a replanning policy; over-planning trivial tasks wastes tokens.
ReWOO: reasoning without observation
ReWOO goes further. The planner writes the whole plan with variable placeholders for tool outputs (e.g. #E1 = Search[x]; #E2 = LLM[summarize #E1]), workers fill in the evidence, and a solver composes the answer. Because the planner never re-reads intermediate observations, it avoids resending the growing history. The paper reports about 5× token efficiency and a 4% accuracy gain on HotpotQA. Xu+ 2023 The cost is zero adaptivity: if #E1 comes back empty, the plan doesn't react.
LLMCompiler: plans as parallel DAGs
LLMCompiler borrows from classical compilers. A Function Calling Planner emits a DAG of tool calls with dependencies, a Task Fetching Unit dispatches calls as soon as their inputs are ready, and an Executor runs independent calls in parallel. Against ReAct it reports up to 3.7× latency speedup, 6.7× cost savings and ~9% accuracy improvement. Kim+ 2023 It also supports replanning when results require it. Modern APIs' parallel tool calls (several tool calls in one model turn) are a lightweight, model-native version of the same idea, limited to one "layer" of the DAG per turn.
Reflexion and self-critique
Reflexion lets an agent learn across trials without weight updates. After a failed attempt, it writes a verbal reflection ("I failed because I didn't check edge case X"), stores it in an episodic memory buffer, and conditions the next attempt on it. It reported 91% pass@1 on HumanEval against 80% for the GPT-4 baseline at the time. Shinn+ 2023 Self-Refine is the single-trial version: generate → self-feedback → refine. Madaan+ 2023
Assuming self-critique always helps. It works best when the critique is grounded in an external signal (test failures, compiler errors, a retrieval check, a stronger or differently-prompted judge). Pure introspection by the same model on the same context often just rubber-stamps or randomly changes correct answers. In design answers, say where the feedback signal comes from.
Tree search: Tree of Thoughts and LATS
Tree of Thoughts (ToT) generalizes CoT into search over "thoughts" (coherent intermediate steps). The model proposes several candidates per step, evaluates them (value prompt or vote), and runs BFS/DFS with backtracking. On Game of 24, GPT-4 with CoT solved 4% of tasks versus 74% with ToT. Yao+ 2023 LATS (Language Agent Tree Search) brings Monte Carlo Tree Search to acting agents: LM-generated value estimates, self-reflections, and real environment feedback guide expansion. It reported 92.7% pass@1 on HumanEval with GPT-4. Zhou+ 2023
The catch is cost: search multiplies LLM calls by the branching factor × depth, often 10–100× a single ReAct run. It also needs either reversible environments (you can't "backtrack" a sent email) or simulation. In 2026 practice, explicit tree search is rare in products. Its role has largely moved into reasoning models' internal thinking and into best-of-n with a verifier (e.g. run n coding attempts in parallel sandboxes, pick the one that passes tests).
CodeAct: code as the action space
Instead of one JSON tool call per turn, the agent writes executable Python that can call multiple tools, loop, branch, and store intermediate values in variables. The interpreter's output (including tracebacks) is the observation. Across 17 LLMs, CodeAct outperformed JSON/text action formats with up to 20% higher success rate. Wang+ 2024 Hugging Face's smolagents makes CodeAgent its first-class agent type, with sandboxed execution via E2B, Modal, Docker and others. HF smolagents docs
# One CodeAct step replaces ~N JSON tool-call turns:
results = [search(f"{city} population 2025") for city in ["Tokyo", "Delhi", "Shanghai"]]
pops = {c: parse_number(r) for c, r in zip(["Tokyo", "Delhi", "Shanghai"], results)}
print(max(pops, key=pops.get), pops) # only this printed summary enters context
Why it works: code is composable (loops and function nesting), models have seen huge amounts of it in pretraining, and intermediate data stays in interpreter variables instead of flowing through the context window. Costs: you need a real sandbox (arbitrary code execution is a security boundary), and failures are harder to attribute than with typed tool calls. The same idea appears as "code execution with MCP" or programmatic tool calling in vendor APIs, where the model writes code that calls tools so large intermediate results never enter context.
| Pattern | Reason ↔ act timing | LLM calls | Adaptivity | Best for | Main risk |
|---|---|---|---|---|---|
| ReAct | Interleaved, every step | 1 per action | High | Default; unpredictable environments | Wandering, loops, quadratic context |
| Plan-and-execute | Plan first, replan on events | Few big + many small | Medium | Long multi-step tasks; human plan approval | Stale plans |
| ReWOO | All reasoning up front | ~2 (plan + solve) | Low | Predictable tool chains; token savings | Can't react to surprises |
| LLMCompiler | DAG plan, parallel exec | Plan + join (+ replan) | Medium | Many independent lookups; latency-sensitive | Wrong dependency graph |
| Reflexion / evaluator loop | After an attempt | × trials | Across trials | Tasks with a checkable outcome (tests) | Ungrounded self-critique |
| ToT / LATS | Search over branches | × branching × depth | High (backtracking) | Puzzles, code with verifiers, offline | Cost; irreversible actions |
| CodeAct | Interleaved, but multi-op per step | Fewer turns | High | Data wrangling, multi-tool composition | Sandbox security; debuggability |
"ReAct vs plan-and-execute?" A strong answer frames it as a latency/cost vs adaptivity trade-off on a known axis. ReAct re-decides after every observation (adaptive, expensive). Plan-first saves calls and gives an approvable artifact but needs a replanning trigger. The realistic answer is a hybrid: plan coarsely, execute with a ReAct sub-loop, replan on failure or surprise, and parallelize independent steps (LLMCompiler-style). Bonus: mention that reasoning models blur the line because they plan internally between tool calls.
- ReAct (Yao+ 2022): the foundational interleaving paper; short and readable.
- LLMCompiler (Kim+ 2023): the DAG-parallel planning design, with latency/cost analysis.
- CodeAct (Wang+ 2024): the case for code as a unified action space.
- LATS (Zhou+ 2023): MCTS over agent actions; good for understanding search costs.
5. Planning in depth
"Planning" in agents covers three separate decisions: when to plan (upfront vs interleaved), how to represent the plan, and how deep to decompose (flat vs hierarchical).
Upfront vs interleaved replanning
Upfront plan
Produce the whole plan before acting. It's cheap, inspectable and parallelizable (a DAG), and humans can approve it. It works when the environment is predictable and early steps don't change what later steps should be. It fails when reality diverges: a missing file, an API returning nothing, a wrong assumption about the codebase.
Interleaved / replanning
Plan a bit, act, observe, revise. Triggers for replanning include a step failing, an observation contradicting a plan assumption, a periodic checkpoint (every k steps), or new user input. It costs more, but it's robust. Most effective coding and research agents keep a high-level plan and revise it continually.
Plan representations
| Representation | What it looks like | Pros | Cons | Seen in |
|---|---|---|---|---|
| Free-text plan in context | "1. find config 2. patch 3. test" | Zero infra | Gets lost as context grows; not machine-checkable | Plan-and-Solve prompts |
| Structured todo list (tool-managed) | todo_write([{id, content, status}]) | Model re-reads and updates it; progress visible to UI and user; fights goal drift | Still flat; model can mark items done falsely | Coding agents (e.g. Claude Code's todo tool) |
| DAG of tasks | Nodes = tool calls/subtasks, edges = data deps | Parallel execution; explicit dependencies | Harder for the model to produce correctly; rigid | LLMCompiler, ReWOO, HuggingGPT |
| Plan/progress files on disk | PLAN.md, progress.txt, features.json + git history | Survives context resets and sessions; humans can edit; other agents can read | Needs discipline to keep it updated | Long-running harnesses Anthropic 2025 |
| Hierarchical task network | Goal → subgoals → primitive actions | Scales to big tasks; maps onto multi-agent delegation | Decomposition errors propagate down the tree | Hierarchical multi-agent, HuggingGPT-style task planning Shen+ 2023 |
Anthropic's long-running-agent harness (Nov 2025) is a good concrete example of "plan as files". An initializer agent runs once to set up the environment: an init.sh, a claude-progress.txt log, an initial git commit, and a JSON feature list where every feature starts with a "passes" status of false. Each later coding agent session reads the progress file and git log, picks one feature, implements it, verifies it end to end (browser automation), commits, and updates progress. Anthropic 2025 The plan lives outside any single context window, so the work survives context resets.
Hierarchical decomposition
For big tasks, decompose recursively until the leaves are "one agent, one context window, verifiable". Rules of thumb:
- Decompose by independence, not by noun. Good subtasks have minimal shared state and a clear output contract.
- Each leaf needs a done-criterion the executor can check (test passes, N sources found, schema-valid output).
- Depth costs context: each level of delegation is a lossy compression of intent. Two levels is common; more than three rarely pays off.
When does planning help?
- Helps: long-horizon tasks (more than about 10 steps), decomposable tasks with parallelizable parts, tasks where a human should approve the approach, and tasks where the agent must manage progress across context windows.
- Hurts or wastes effort: short tasks (planning overhead is larger than the task), highly reactive environments (plans go stale immediately), and tasks where the model's first plan is confidently wrong and nothing triggers a replan.
- 2026 nuance: reasoning models do a lot of implicit planning in their thinking tokens, so explicit planning scaffolds give smaller gains on short tasks than they did in 2023. The persistent, externalized plan (todo/progress file) is still clearly valuable for long tasks, mainly as a memory and anti-drift device.
"How does your agent plan?" Interviewers want to hear about plan representation and lifecycle, not just "it makes a plan". A strong answer: a structured todo list the agent re-reads every turn, explicit replan triggers, done-criteria per item, and persistence outside the context window for long tasks. Plus awareness that the agent may mark items done that aren't, which is why done is verified by the environment, not by the model's say-so.
- Anthropic: Effective harnesses for long-running agents: initializer agent, progress files, one-feature-at-a-time.
- Plan-and-Solve prompting (Wang+ 2023): the minimal "plan then solve" result.
- HuggingGPT (Shen+ 2023): early LLM-as-controller doing task planning → model selection → execution.
6. Multi-agent architectures
A "multi-agent system" is any system where more than one LLM loop, each with its own prompt, tools and (usually) its own context window, cooperates on a task. The reasons to split are always some mix of: context isolation (each agent sees only what it needs), parallelism (wall-clock speed, more total tokens spent on the problem), specialization (different prompts/tools/models per role), and organizational modularity (teams own different agents). The costs are always: lost context at boundaries, coordination overhead, more tokens, and harder debugging.
Think of agents as microservices whose API is natural language. All the distributed-systems lessons apply: chatty interfaces are slow, shared mutable state needs coordination, the network (here, the context handoff) is lossy, and "we split it into services" doesn't fix a bad domain model. The extra wrinkle is that each service is also non-deterministic and confidently fills gaps in its spec with guesses.
(a) Orchestrator–worker / lead agent with subagents
A lead agent plans, spawns subagents with specific subtasks (often in parallel), receives their condensed results, and synthesizes. Subagents usually can't talk to each other. Everything goes through the lead.
The reference implementation is Anthropic's Research feature (June 2025). A lead agent (Claude Opus 4) plans, saves the plan to memory (in case the context gets truncated), spawns subagents (Claude Sonnet 4) that search in parallel, then synthesizes, with a separate CitationAgent pass to attribute sources. Anthropic 2025 The reported numbers, worth memorizing with their caveats:
- The multi-agent system outperformed single-agent Opus 4 by 90.2% on Anthropic's internal research eval. This is an internal benchmark of breadth-first queries, not a general result.
- Token usage alone explained ~80% of performance variance on BrowseComp, with tool-call count and model choice making up the rest of the roughly 95% explained. This is the key mechanistic insight: much of multi-agent's advantage is that it spends more tokens in parallel, across more context windows than one agent could hold.
- Agents used about 4× the tokens of chat interactions, and multi-agent systems about 15×.
- The lead typically spawned 3–5 subagents in parallel, and subagents used 3+ tools in parallel. Together these cut research time by up to 90% for complex queries.
Engineering lessons from that post that come up in interviews: teach the orchestrator to scale effort to query complexity (simple fact-finding gets 1 agent with a few tool calls; complex research gets many subagents); give subagents detailed task descriptions (objective, output format, tool guidance, boundaries), because vague delegation caused duplicated work and gaps; use "start wide, then narrow" search strategies; and have subagents write outputs to a filesystem/artifact store and pass back lightweight references, so large content doesn't get squeezed through the lead's context. Anthropic 2025
Pros
- Parallel breadth: wall-clock speed and more total tokens on the problem
- Context isolation: each subagent's window is clean and focused; the lead stays small
- Natural fit for "read-heavy" tasks (research, search, code exploration, review)
- The lead keeps one coherent decision-making thread
Cons
- About 15× chat tokens; only worth it for high-value tasks
- The lead is a bottleneck and single point of failure; synchronous waves block on the slowest subagent
- Subagents lack each other's context, so you get duplicated work or inconsistent assumptions
- Lossy compression of findings back to the lead
(b) Supervisor / router
A supervisor agent decides, at each turn, which specialist agent acts next (or routes the whole request once). Specialists return control to the supervisor after they act. It differs from orchestrator–worker mostly in emphasis: the supervisor is a dispatcher over a fixed roster of specialists (often sequential), whereas the orchestrator creates subtasks dynamically (often parallel). In OpenAI's Agents SDK terms this is the "agents as tools" / manager pattern: the manager calls specialists as tools and keeps ownership of the final answer. OpenAI docs
- Pros: central control and a single place for policy; easy to add specialists; one coherent voice to the user.
- Cons: an extra LLM hop per turn (latency, cost); the supervisor's routing quality caps the system; if specialists return verbose outputs, the supervisor's context bloats.
(c) Hierarchical teams
Supervisors of supervisors: a top-level coordinator delegates to team leads (e.g. "research team", "writing team"), each managing their own workers. It mirrors org charts and is how MetaGPT-style "software company" simulations are structured, with roles like product manager, architect and engineer following SOPs. Hong+ 2023
- Pros: scales to large task trees; each supervisor's span of control stays small; teams can be developed and evaluated separately.
- Cons: every level is a lossy re-statement of intent (the telephone game); latency multiplies with depth; hard to trace failures to a level. Rarely justified beyond two levels.
(d) Handoffs / swarm
Decentralized transfer of control: the active agent decides to hand the conversation to another agent, which then becomes the active agent and talks to the user directly. There's no central supervisor in the loop. OpenAI's experimental Swarm library introduced this pattern, and its README now says it's replaced by the OpenAI Agents SDK, "a production-ready evolution of Swarm". OpenAI Swarm In the Agents SDK, handoffs are exposed to the model as tools: a handoff to an agent named "Refund Agent" appears as a tool named transfer_to_refund_agent. By default the receiving agent sees the full conversation history, and an input_filter can trim what it sees. OpenAI Agents SDK docs
- Pros: no extra supervisor hop (lower latency); each specialist has narrow instructions and tools; natural for customer-service triage (triage → billing → refunds).
- Cons: no global view, so agents can ping-pong; harder to enforce cross-cutting policy; the full history passed along can confuse the receiver or leak irrelevant context; poor at multi-domain requests that need synthesis. LangChain's comparison notes handoffs are efficient for single-domain and repeat requests but weak on multi-domain queries. LangChain docs
(e) Sequential pipelines / assembly line
Fixed stages, each an agent (or LLM call) with its own role: researcher → writer → editor → fact-checker, or localize → repair → validate for code. This is really prompt chaining where the stages happen to be agentic. Agentless is a well-known example: a fixed three-phase localization → repair → patch-validation process with no LLM-chosen control flow. It scored 32% on SWE-bench Lite at about $0.70 per issue, beating the open-source agents of the time. Xia+ 2024
- Pros: predictable, testable per stage, easy to observe and cost-bound; each stage can use a different model.
- Cons: no backtracking unless you add explicit feedback edges; errors in early stages propagate; rigid for inputs that don't fit the pipeline.
(f) Debate / multi-agent critique
Several agents (same or different models, possibly with assigned stances) propose answers, read each other's reasoning, and revise over rounds until they converge or a judge decides. Du et al. showed that multi-agent debate improved math and strategic reasoning and reduced factual errors. Du+ 2023 The generator–verifier loop is the practical cousin: a reviewer agent with a clean context checks the producer's output.
- Pros: catches errors a single pass misses, especially when the critic has fresh context or is a different model family (decorrelated errors).
- Cons: costs rounds × agents × tokens; agents tend to converge on a confident wrong consensus (conformity), so the gain over plain self-consistency voting is contested; needs a stopping rule.
(g) Blackboard / shared state
A classic AI architecture: agents don't message each other. They read from and write to a shared workspace (the "blackboard"), and a controller (or the agents themselves) decides who acts next based on the blackboard's state. In LLM systems the blackboard is a shared document, a structured state object (LangGraph's graph state is effectively this), a database, or a repo.
- Pros: loose coupling (agents only know the shared schema); everything is inspectable; agents can be added without rewiring messages; natural for incremental, opportunistic problem solving.
- Cons: concurrent writes need conflict resolution (locks, CRDTs, versioned updates, reducers); the blackboard grows and needs curation; agents can overwrite each other's work, which is exactly Cognition's "conflicting implicit decisions" problem.
(h) Market / auction, peer-to-peer and network topologies
Market/auction (contract-net): a manager announces a task, agents bid (with estimated cost, confidence or capability), and the best bid wins. It's useful when agents are heterogeneous (different costs/tools) and the roster is dynamic, e.g. cross-organization agents discovered via A2A agent cards. Peer-to-peer / network: any agent can message any other (AutoGen-style group chats, CAMEL role-playing Li+ 2023). It's flexible, but communication grows as \(O(n^2)\), it's hard to reason about termination, and it's mostly a research or simulation setup. Generative Agents (the "Smallville" simulation) is the famous example of agent societies used for simulation rather than task completion. Park+ 2023
As of 2026, practitioner consensus (including Cognition's April 2026 follow-up) is that unstructured swarms, meaning arbitrary networks of agents negotiating with each other, mostly don't see meaningful production adoption, while structured manager/verifier patterns do. Cognition 2026 Market-based coordination among LLM agents is still largely research. Check the current literature before presenting it as production practice.
Summary table
| Topology | Control | Communication | Parallelism | Best for | Biggest risk |
|---|---|---|---|---|---|
| Orchestrator–workers | Central, dynamic decomposition | Task spec down, summary up | High | Breadth-first research, codebase exploration, review | Subagents lack shared context; 15× tokens |
| Supervisor / router | Central dispatcher | Hub and spoke | Low–medium | Fixed set of specialists; policy centralization | Extra hop latency; supervisor bottleneck |
| Hierarchical | Tree of supervisors | Up/down the tree | Medium–high | Very large decomposable tasks | Intent loss per level |
| Handoffs / swarm | Decentralized; active agent transfers | Shared conversation passed along | None (one active) | Conversational triage, customer support | Ping-pong; no global view |
| Sequential pipeline | Fixed by code | Stage outputs | None (but can pipeline across items) | Repeatable processes, content pipelines | Early errors propagate; rigid |
| Debate / critique | Rounds + judge | Broadcast of proposals | High per round | High-stakes answers, review, verification | Cost; conformity |
| Blackboard | Controller or opportunistic | Shared state, no direct messages | Medium | Incremental problem solving; many contributors | Write conflicts; state bloat |
| Market / P2P | Emergent / bidding | Any-to-any | High | Heterogeneous cross-org agents; simulations | Termination, \(O(n^2)\) chatter, unpredictability |
"Supervisor vs handoff?" Ask who should own the final answer. If one agent must synthesize across specialists and keep consistent policy, use a supervisor or agents-as-tools. If a specialist should take over the conversation for a branch (e.g. refunds) and talk to the user directly, use a handoff. OpenAI's own docs frame the choice exactly this way, and also say "start with one agent whenever you can". OpenAI docs
- Anthropic: How we built our multi-agent research system: the most detailed public account of a production orchestrator–worker system.
- OpenAI Agents SDK: orchestrating multiple agents: LLM-driven vs code-driven orchestration, handoffs vs agents-as-tools.
- Magentic-One (Fourney+ 2024): Microsoft's generalist orchestrator with task/progress ledgers and specialist agents (web, file, coder, terminal).
- AutoGen (Wu+ 2023): the multi-agent conversation framework paper that popularized group-chat topologies.
7. Agent roles and how they differ
Role names get used loosely in interviews and papers. Here's a precise vocabulary. One agent can play several roles, and a role can be plain code instead of an LLM.
| Role | Responsibility | Owns | Often implemented as | Don't confuse with |
|---|---|---|---|---|
| Planner | Turns a goal into steps / a DAG; replans on events | The plan artifact | Strong reasoning model, called rarely | Orchestrator (planner doesn't execute or delegate) |
| Orchestrator | Decomposes at runtime, spawns workers, integrates results | Task lifecycle + synthesis | Lead agent with a "spawn subagent" tool | Router (one-shot dispatch, no synthesis) |
| Coordinator / supervisor | Chooses which existing agent acts next; enforces turn-taking and policy | Control flow among a fixed roster | LLM with "route_to(agent)" tool, or code | Orchestrator (creates new subtasks) |
| Router | Classifies input once and dispatches | A single routing decision | Small model / classifier / embeddings | Supervisor (multi-turn control) |
| Executor / worker | Performs a bounded subtask with tools | Its subtask's output | ReAct loop with narrow tools; cheaper model | Planner |
| Critic / verifier | Checks outputs against criteria; returns pass/fail + feedback | Quality gate | Judge LLM with clean context, tests, linters, schema validators | Generator self-review (same context, correlated errors) |
| Memory manager | Decides what to persist, summarize, retrieve, forget | Long-term store + compaction | Background LLM job or code policy (see Memory) | Blackboard (shared working state, not long-term memory) |
| Synthesizer / reporter | Merges worker outputs into a final artifact | Final answer | Often the orchestrator itself | Citation agent (post-hoc attribution) |
Making every role a separate LLM agent. A verifier that's really "run the test suite" should be code. A router over 4 well-separated intents can be a fine-tuned classifier at about 1/100th the latency. Use an LLM for a role only when the role needs judgment.
8. The case against vs the case for multi-agent
This was the defining architecture debate of mid-2025, and interviewers love it because it tests judgment rather than recall. The two key posts came out one day apart in June 2025.
Against: Cognition, "Don't Build Multi-Agents" (Walden Yan, 12 Jun 2025)
Two principles Yan 2025:
- Share context, and share full agent traces, not just individual messages. A subagent given only a task description lacks the decisions and nuance in the parent's history.
- Actions carry implicit decisions, and conflicting decisions carry bad results. Parallel agents make unstated assumptions (style, interfaces, approach) that clash when combined.
Example: "build a Flappy Bird clone" split into "background" and "bird" subtasks. One subagent builds a Super Mario-style background, the other a bird that doesn't match, and the combiner has to reconcile incompatible work. Recommendation: a single-threaded linear agent, and for very long tasks a dedicated model that compresses history into key details and decisions.
For: Anthropic, multi-agent research system (13 Jun 2025)
For breadth-first research queries with many independent directions, parallel subagents with isolated contexts won big (+90.2% on the internal eval). Token spend explains most of the variance, and multi-agent is a way to spend more tokens usefully by spreading them across many context windows in parallel. Anthropic 2025
The same post is explicit about limits: multi-agent fits less well for domains that need all agents to share the same context or that have many dependencies between agents, and it notes that most coding tasks have fewer truly parallelizable parts than research. It's also expensive (about 15× chat tokens), so the task's value has to justify it.
The synthesis
Read the two posts together and they mostly agree. LangChain's commentary at the time framed it as reading vs writing: multiple agents gathering information (reading) parallelize well, while multiple agents producing parts of one artifact (writing) create conflicting implicit decisions that are hard to merge. Chase 2025 Cognition's own April 2026 follow-up, "Multi-Agents: What's Actually Working", lands on the same principle: writes stay single-threaded, and additional agents contribute intelligence rather than actions. The patterns it reports working are a clean-context code-review agent (generator–verifier), a "smart friend" frontier model consulted by the primary agent, and manager agents that coordinate scoped child agents. Parallel-writer swarms still don't work well. Cognition 2026
Controlled research points the same way. A Dec 2025 study by Google researchers and collaborators evaluated 260 configurations across six agentic benchmarks and five architectures (single-agent; independent, centralized, decentralized and hybrid multi-agent). Multi-agent performance relative to single-agent ranged from +80.8% on decomposable financial reasoning to −70.0% on sequential planning. Coordination showed diminishing returns once the single-agent baseline was already strong, tool-heavy tasks paid a multi-agent overhead, and architectures without centralized verification propagated errors more. Kim+ 2025 The MAST failure taxonomy (1,600+ annotated traces across 7 multi-agent frameworks) groups failures into system design issues, inter-agent misalignment, and task verification, and notes that multi-agent gains on popular benchmarks are often minimal. Cemri+ 2025
| Multi-agent tends to win when… | Single agent tends to win when… |
|---|---|
| The task is breadth-first / embarrassingly parallel (many independent sub-questions, many files to read) | The task is sequential / tightly coupled (each step depends on the last; one coherent artifact) |
| Subtasks are read-only or produce independent outputs | Subtasks write to shared artifacts with implicit design decisions |
| Total information exceeds one context window, and subagents can compress it | Everything fits in one context (perhaps with compaction) |
| Wall-clock latency matters and parallel spend is affordable | Cost per task matters more than latency |
| You want an independent verifier with clean context | Coordination overhead would dominate (simple tool-heavy tasks) |
| Task value is high (research reports, large migrations) | High-volume, low-value requests (support FAQs) |
If asked "single or multi-agent?", don't pick a side. A strong answer goes: "It depends on task structure. I'd parallelize the read-heavy, independent parts (research, exploration, review) with isolated subagents returning condensed results, and keep any writes to a shared artifact in a single thread that owns the decisions. Expect about 15× tokens versus chat, so only for high-value tasks, and I'd validate the gain on my own evals against a strong single-agent baseline with compaction." Name both posts and the Cognition 2026 update.
- Cognition: Don't Build Multi-Agents (2025): the context-sharing argument with the Flappy Bird example.
- Cognition: Multi-Agents: What's Actually Working (Apr 2026): the follow-up; verifier, smart-friend and manager patterns.
- Towards a Science of Scaling Agent Systems (Kim+ 2025): controlled comparison of single- vs multi-agent architectures by task type.
- Why Do Multi-Agent LLM Systems Fail? (Cemri+ 2025): the MAST taxonomy of 14 failure modes.
9. Inter-agent communication and handoff context
Agents communicate in three basic ways, and the choice decides what context survives a boundary.
| Mechanism | How | Pros | Cons | Example |
|---|---|---|---|---|
| Shared transcript | All agents read/append one message history | Maximum shared context; simple | Context bloat; agents distracted by irrelevant turns; prompt confusion about "who am I" | Group chat; Agents SDK handoffs (full history by default) |
| Message passing | Agent A sends a structured message/task to B; B returns a result | Isolation; small contexts; clear contracts | Lossy: B only knows what A wrote down | Subagent-as-tool; A2A tasks |
| Artifacts / files / shared state | Agents write to files, DB, state object; pass references | Large content doesn't squeeze through contexts; persistent; inspectable | Needs schemas, versioning, conflict handling | Research subagents writing to a filesystem; LangGraph state; repos |
What to pass in a handoff
Most multi-agent failures are really specification failures at the boundary. Every subagent starts with zero knowledge except its prompt. A good structured task spec looks like this:
{
"objective": "Find 2024-2026 benchmark results for open-weight models on SWE-bench Verified",
"why": "Parent is writing a comparison table for a CTO memo; needs numbers + sources",
"scope": {"include": ["official leaderboards", "model cards"], "exclude": ["blog speculation"]},
"known_context": ["Parent already has results for closed models; don't repeat them"],
"constraints": {"max_tool_calls": 15, "max_tokens_out": 1200, "deadline_s": 120},
"output_format": {"type": "json", "schema": "[{model, score, date, source_url}]"},
"done_when": ">= 5 models with primary-source URLs, or budget exhausted (report gaps)",
"decisions_already_made": ["Use 'Verified' split only", "Report pass@1"]
}
Notice why and decisions_already_made. They're the direct mitigation for Cognition's "implicit decisions" problem: you write down the decisions the subagent would otherwise make differently. Also ask for condensed outputs with references (findings + source pointers, or a file path) rather than raw transcripts, so the parent's context stays clean.
Returning results
- Summaries, not transcripts. A subagent might burn 50k tokens exploring but should return 1–2k tokens of distilled findings. That compression is the context-isolation benefit.
- Include uncertainty and gaps: "couldn't find X", "sources disagree on Y". Otherwise the parent assumes completeness.
- Structured outputs (JSON schema) make synthesis and validation mechanical.
- Large artifacts by reference: write to a file or object store and return the path plus a short summary.
Handoffs that pass the full conversation to a narrow specialist. The specialist's instructions get diluted by thousands of irrelevant tokens and earlier agents' persona text. Filter history at handoff time: the Agents SDK exposes input_filter for exactly this. OpenAI Agents SDK docs The opposite failure, passing a one-line task, is just as common. The right amount is a structured spec plus relevant excerpts.
10. State management and control flow
Once an agent does more than one tool loop, you need to answer: where does state live, who decides the next step, and what happens if the process dies or a human must approve something? The dominant answer since 2024 is to model the agent as an explicit state graph.
Graph-based state machines (LangGraph-style)
In LangGraph, you define a typed state object, nodes (functions, which may call LLMs or tools, that read state and return updates), and edges (fixed or conditional, where a function of state picks the next node). Updates are merged into state via per-key reducers (e.g. append to the message list rather than overwrite), which is what makes parallel branches safe to join. Execution proceeds in "super-steps", and with a checkpointer attached a state snapshot is saved at every super-step, keyed by a thread_id. LangGraph docs The same model covers fixed workflows (all edges fixed), ReAct agents (an LLM node ↔ tool node cycle with a conditional "tool calls?" edge), and multi-agent systems (nodes are subgraphs).
Checkpoints, human-in-the-loop and resumability
- Checkpoints give you: resume after a crash or deploy, "time travel" (inspect or fork from an earlier state to debug), multi-turn conversations (same
thread_id), and HITL pauses that can last days. In production, use a persistent checkpointer such as Postgres. - Interrupts: a node calls
interrupt(payload), the graph saves state and returns, and later you resume withCommand(resume=value), which becomes the return value ofinterrupt. One subtlety worth knowing: on resume, the whole node re-executes from its beginning, so side effects placed before theinterruptcall run again and should be idempotent or moved after it or into a separate node. LangGraph docs - HITL patterns: approve/reject a tool call; edit the tool arguments; edit the state (fix the plan); answer a clarifying question; review the final output. Put approvals on irreversible or high-blast-radius actions (payments, emails, deletes, prod deploys), not on everything, because approval fatigue turns humans into rubber stamps.
Treating "checkpointed state" as "durable execution". A checkpointer saves state between steps. It doesn't by itself guarantee that a step interrupted halfway through a non-idempotent side effect is handled correctly, or that a crashed process gets rescheduled automatically. For that you need a durable execution engine or your own retry and idempotency discipline (next section).
"How would you add human approval to an agent that can issue refunds?" Strong answer: model the refund as a tool whose execution routes through an approval node; persist state at the interrupt (checkpointer keyed by conversation/thread); notify a reviewer with the proposed arguments plus the agent's rationale; resume with approve/edit/reject; make the refund call idempotent with a key derived from (ticket, amount) so a resume or retry can't double-pay; set a timeout path if no one approves. Bonus: policy thresholds, e.g. auto-approve under a set amount.
- LangGraph: Persistence: checkpoints, threads, super-steps, time travel.
- LangGraph: Interrupts: HITL semantics, including the node re-execution gotcha.
- LangGraph: Durable execution: persistence modes and how to wrap side effects.
11. Durable execution for long-running agents
An agent that runs for minutes to hours, waits on humans for days, or calls flaky external APIs is a long-running distributed workflow, and it should be built like one. The lessons come straight from workflow engines like Temporal.
The Temporal model, applied to agents
Temporal splits code into workflows (deterministic orchestration code whose progress is recorded in an event history and recovered by replay after a crash) and activities (the non-deterministic, side-effecting work: network calls, DB writes, and, for agents, LLM calls and tool executions), which get automatic retries with configurable policies. Temporal docs Temporal docs Mapped onto an agent:
Why this matters: an LLM response is non-deterministic, so it must be an activity whose result is recorded. Otherwise replay would get a different answer and diverge. The workflow keeps only control logic. Temporal has published integrations with agent frameworks (for example a LangGraph plugin) that push LLM and tool calls into activities. Temporal blog
Idempotent tools, retries and sagas
- Idempotency keys: any side-effecting tool takes a key (e.g.
hash(run_id, step_id)) so a retry after a timeout ("did the email send?") doesn't duplicate. Prefer upserts to inserts. - Retry classes: retry transient errors (429, 5xx, timeouts) with exponential backoff and jitter. Don't blindly retry semantic errors (400, validation). Return those to the model as observations so it can fix its arguments.
- Sagas / compensation: for multi-step side effects across systems (book flight → book hotel → charge card), record a compensating action for each completed step (cancel flight, cancel hotel) and run them in reverse order if a later step fails. microservices.io For agents, declare compensations in the tool registry, because an LLM shouldn't improvise rollback.
- Timeouts at every level: per tool call, per LLM call, per step, per run. Hung tools are a top cause of stuck agents.
- Versioning: long-running runs outlive deploys. Workflow code changes need versioning so in-flight runs replay correctly, and prompts and tool schemas should be versioned too.
"Your agent runs a 2-hour data migration and the pod gets killed at minute 70. What happens?" A weak answer: "it restarts". A strong answer: state is checkpointed per step (or the run is a durable workflow); completed steps' LLM and tool results are recorded and replayed, not re-executed; the in-flight step's side effect is idempotent (keyed) so re-running it is safe; non-reversible steps were gated; and partial failure triggers compensations. Then mention observability: a run ID tying traces, cost and state together.
- Temporal: Workflows: deterministic replay, event history, signals.
- Temporal: Retry policies: backoff, non-retryable errors.
- Saga pattern (microservices.io): compensating transactions, choreography vs orchestration.
12. Concurrency and parallel tool calls
There are three levels of parallelism in agent systems:
- Parallel tool calls within one turn. Modern model APIs let the model emit several tool calls in one response (e.g. three searches). The harness runs them concurrently (
asyncio.gather) and returns all results together. It's cheap to add and cuts latency for independent lookups. In Anthropic's research system, parallel tool calls in subagents were one of the two parallelization levers behind the "up to 90%" time reduction. Anthropic 2025 - Parallel branches in a workflow / graph. Fan out to nodes, fan in with reducers (LangGraph), or use a DAG executor (LLMCompiler).
- Parallel agents. Subagents in separate contexts, possibly separate processes or sandboxes.
| Concern | What goes wrong | Mitigation |
|---|---|---|
| Hidden dependencies | Model parallelizes calls that depend on each other (edit file, then run test, concurrently) | Mark tools as read-only vs mutating; serialize mutating calls; prompt guidance |
| Write conflicts | Two agents edit the same file / record | Single writer; locks; git worktrees per agent plus merge; optimistic concurrency with versions |
| Rate limits | Fan-out of 20 agents × 5 tools hits provider 429s | Semaphores / token buckets per provider; backoff; queue |
| Straggler latency | Synchronous wave waits for slowest subagent | Per-subagent timeouts; accept partial results; async/streaming orchestration |
| Cost amplification | Parallelism multiplies spend quickly | Budget per run split across children; cap fan-out width |
| Ordering in context | Results returned out of order confuse attribution | Match results to call IDs (APIs require tool_use_id / call_id) |
Back-of-envelope: a research step with 5 independent searches at about 2 s each takes about 10 s serially and about 2 s in parallel. Five subagents that each run 10 sequential steps of about 6 s (LLM + tool) finish in about 60 s wall-clock instead of about 300 s, at roughly the same or higher token cost. Parallelism buys latency, not cost.
13. Specialized agent types
Coding agents
Coding is the most commercially mature agent domain, because the environment gives dense, objective feedback (compilers, linters, tests) and the work is reversible (git). The typical loop: understand the task → explore the repo (search, read files) → plan → edit → run tests/linters → read failures → iterate → present the diff.
- Agent–computer interface (ACI). SWE-agent's key idea: LM agents are a new kind of user and need purpose-built interfaces. Its custom ACI had a file viewer showing a window of about 100 lines, search commands with concise output, and an editor that runs a linter and rejects syntactically invalid edits. That substantially improved performance over raw shell access, giving 12.5% pass@1 on SWE-bench at the time, which was state of the art. Yang+ 2024 The general lesson: tool design is prompt design. Anthropic reports spending more time optimizing tools than the overall prompt for its SWE-bench agent. Anthropic 2024
- Edit formats. There are several options: whole-file rewrite (simple, wasteful, risks truncation), search/replace blocks (exact-match old string → new string; the most common in production agents; fails loudly when the match isn't unique), unified diffs (compact but models make line-number errors), and AST- or LSP-aware edits. Requiring a unique exact match is a mistake-proofing trick: it forces the model to read before writing.
- Sandboxes. Run code in containers or microVMs (Docker, Firecracker-based services) with no prod credentials, restricted network egress, and resource limits. Permission tiers decide what's auto-allowed (read files, run tests), what needs approval (network, installing packages, git push), and what's denied.
- Test feedback. Tests are the verifier. Strong agents write a reproduction test first, then fix. Watch for agents that "make tests pass" by editing or deleting the tests, and guard against it with instructions and diff checks.
- Benchmarks: SWE-bench (real GitHub issues verified by hidden tests) Jimenez+ 2023 and its Verified subset are the standard references. Scores have risen fast, so quote current leaderboards, not memorized numbers.
Coding-agent capability, SWE-bench scores, and practices (background/async agents, many parallel agents in git worktrees, AGENTS.md-style repo instructions) changed month to month through 2025–2026. Check swebench.com and vendor posts for current numbers before quoting any.
Computer-use and browser agents
The action space is the GUI: screenshots (plus optionally the accessibility tree or DOM) come in as observations, and mouse/keyboard actions go out. Anthropic, OpenAI and Google all ship computer-use or browser-use model capabilities. Claude computer use docs Key engineering points:
- Grounding (clicking the right pixel/element) is a core failure source. DOM/accessibility-tree actions are more reliable than raw coordinates when they're available, so hybrid approaches are common.
- Latency and cost: each step sends a screenshot (roughly 1–2k image tokens), and tasks take tens of steps, so these agents are slow compared with API-based tools. Prefer an API or MCP tool whenever one exists and use GUI automation as the fallback.
- Security: web content is untrusted input, so prompt injection via page text is the main threat. Isolate in a VM/container, use allowlisted domains, keep credentials out of reach, and confirm sensitive actions with a human.
- Benchmarks: WebArena (realistic self-hosted websites) Zhou+ 2023 and OSWorld (real desktop OS tasks) Xie+ 2024.
Deep-research agents
The canonical orchestrator–worker use case: plan sub-questions → parallel search/read subagents → iterative deepening ("start wide, then narrow") → synthesis → citation verification. Design concerns: source quality ranking (prefer primary sources over SEO content), deduplication across subagents, claim-level citations checked against the fetched text, explicit "what I couldn't find" sections, and scaling effort to the question (don't spawn 10 agents for a factoid). Anthropic 2025 Evaluation is hard because answers are open-ended, so use LLM-judge rubrics (factual accuracy, citation accuracy, completeness, source quality, tool efficiency) plus human spot checks.
Voice agents
Voice agents run under a hard latency budget. In human conversation the gap between turns is short, and responses much slower than about ~1 s start to feel unnatural. That reshapes the architecture:
- Pipeline vs speech-to-speech: a cascaded STT → LLM → TTS pipeline (modular, any LLM, easy to inspect, more latency) vs native speech-to-speech realtime models (lower latency, better prosody, less control). OpenAI Realtime docs
- Latency budget (cascaded, rough): VAD/endpointing ~200–500 ms + STT finalization ~100–300 ms + LLM time-to-first-token ~200–500 ms + TTS time-to-first-audio ~100–300 ms. Everything streams, and the LLM's first sentence goes to TTS before the rest is generated.
- Turn-taking: endpointing (when has the user finished?), barge-in (the user interrupts, so stop TTS immediately and truncate the assistant message to what was actually spoken), and backchannels.
- Tools under latency: slow tool calls need filler speech ("let me check that…"), asynchronous tools, or prefetching. Keep agent graphs shallow, since multi-agent hops add round trips the caller hears as silence. Handoffs (not supervisors) are common in voice because they avoid an extra LLM hop per turn.
- Frameworks: LiveKit Agents and Pipecat are widely used open-source voice-agent frameworks. LiveKit docs Pipecat
| Agent type | Observation | Action space | Feedback signal | Dominant constraint | Typical architecture |
|---|---|---|---|---|---|
| Coding | Files, test output, errors | Read/search/edit/shell | Tests, compilers (dense, objective) | Correctness; safety of shell | Single-threaded writer + read-only explorer/reviewer subagents |
| Computer-use / browser | Screenshots, DOM | Click/type/scroll | Page state (noisy) | Grounding, latency, injection | Single agent in sandboxed VM; HITL for sensitive steps |
| Deep research | Search results, pages | Search/fetch/read | Weak (judge rubrics) | Breadth, source quality, cost | Orchestrator + parallel subagents + citation pass |
| Voice | Transcribed speech (or audio) | Speak + tools | User reaction | Latency (< ~1 s), turn-taking | Streaming single agent or handoffs; minimal hops |
| Customer support | User messages, CRM data | Lookup/refund/escalate | Resolution, CSAT | Policy compliance, cost/volume | Router + narrow tool agents + HITL on money |
- SWE-agent (Yang+ 2024): the agent–computer interface paper; why interface design changes agent performance.
- OpenHands (Wang+ 2024): an open platform for coding/generalist agents built on CodeAct.
- Anthropic: Raising the bar on SWE-bench Verified: a minimal coding scaffold (bash + edit tool) and the tool-design details that mattered.
14. Frameworks landscape (as of late 2026)
This landscape changes every quarter. Status notes below were checked against official docs in October 2026, but versions, names and consolidations (e.g. AutoGen → Microsoft Agent Framework) keep happening. Check current docs before making claims in an interview.
| Framework | From | Core abstraction | Multi-agent style | Notes / 2026 status |
|---|---|---|---|---|
| LangGraph docs | LangChain | Typed state graph: nodes, conditional edges, reducers | Anything as subgraphs: supervisor, swarm, hierarchical | Low-level and explicit; checkpointers, interrupts, durable execution modes. LangGraph reached a 1.0 release in late 2025. The usual pick when you want explicit control flow. |
| OpenAI Agents SDK docs | OpenAI | Agent (instructions + tools), Runner loop | Handoffs; agents-as-tools | Production successor to Swarm; guardrails, sessions, tracing built in. Python and TypeScript. |
| Claude Agent SDK docs | Anthropic | Claude Code's agent loop as a library | Subagents | Built-in file/shell/web tools, hooks, permissions, sessions, MCP, skills; Python and TypeScript. Formerly the Claude Code SDK. |
| CrewAI docs | CrewAI | Role-based "crews" of agents with tasks; "Flows" for event-driven control | Sequential or hierarchical (manager) processes | High-level, role/goal/backstory framing; quick to prototype. |
| Microsoft Agent Framework docs | Microsoft | Agents + graph-based workflows | Orchestration patterns inherited from AutoGen (sequential, concurrent, group chat, handoff, magentic) | Successor that unifies AutoGen and Semantic Kernel; .NET and Python. Public preview Oct 2025, 1.0 GA reported Apr 2026. MS blog AutoGen is in maintenance mode. |
| AG2 GitHub | Community (AutoGen fork) | Conversable agents, group chat | Conversation-driven | Community continuation of the original AutoGen 0.2 line. |
| Google ADK docs | LlmAgent + workflow agents (Sequential, Parallel, Loop); graph workflows in 2.0 | Agent hierarchies; A2A integration | Python, TypeScript, Go, Java, Kotlin; tight Gemini/Vertex integration but model-agnostic. | |
| LlamaIndex docs | LlamaIndex | Event-driven Workflows; agents over data/RAG | Multi-agent workflows | Strongest on document/data-centric agents and retrieval. |
| Pydantic AI docs | Pydantic | Type-safe Agent with typed deps and structured outputs | Delegation via tools, programmatic handoff, graphs | "FastAPI feel"; strong validation; model-agnostic. |
| smolagents docs | Hugging Face | CodeAgent (code as action) and ToolCallingAgent | Managed agents | Minimal (core logic around a thousand lines); sandboxed execution options. |
| DSPy docs | Stanford NLP | Declarative modules (signatures) + optimizers that compile prompts/few-shots | Compose modules (e.g. dspy.ReAct) | Not an orchestration framework so much as a way to optimize LM programs against a metric. Khattab+ 2023 |
Framework vs roll your own
Use a framework when…
- You need checkpointing, HITL interrupts, streaming, and tracing now, not after a quarter of infra work
- The team is new to agents and benefits from opinionated structure
- You want ecosystem integrations (MCP, vector stores, observability)
- Your control flow is graph-shaped and you'd otherwise reinvent a state machine
Roll your own when…
- The core loop is ~100 lines and you want full visibility into every prompt and token
- You need tight latency (voice) or unusual control flow
- You already have a workflow engine (Temporal, Step Functions) for durability
- Framework abstractions fight your model provider's newest features (parallel tools, caching, thinking)
The common middle path: a thin in-house loop on the provider SDK, a real workflow engine for durability, MCP for tool integration, and OpenTelemetry-style tracing. Whichever you pick, make sure you can see the exact prompt sent to the model at every step. That's the debugging superpower frameworks sometimes hide.
"Which framework would you use?" isn't a trivia question. They want trade-off reasoning. A good answer: "For an explicit, auditable workflow with HITL, something graph-based like LangGraph or a Temporal workflow. For a coding-style agent that needs file and shell tools, the Claude Agent SDK. For OpenAI-centric triage with handoffs, the Agents SDK. But I'd prototype the loop by hand first to understand the prompts, then adopt a framework for persistence and tracing, not for the loop itself."
15. Inter-agent protocols: MCP and A2A
MCP: Model Context Protocol
An open protocol, introduced by Anthropic in Nov 2024, for connecting AI applications to tools, resources and prompts exposed by servers. Anthropic 2024 Client–server over JSON-RPC (local stdio or streamable HTTP). Servers expose tools (callable functions), resources (readable data) and prompts (templates). The point is to solve the M×N integration problem once: any MCP client can use any MCP server. Spec versions are date-stamped, e.g. 2025-11-25. In Dec 2025 Anthropic donated MCP to the Linux Foundation's new Agentic AI Foundation (AAIF), co-founded with Block and OpenAI. MCP blog 2025
A2A: Agent2Agent Protocol
An open standard, created by Google (Apr 2025) and donated to the Linux Foundation, for agent-to-agent communication across frameworks and vendors. Agents interoperate without exposing internal memory, tools or logic. A2A docs Core concepts: an Agent Card (a JSON description of an agent's identity, skills, endpoint and auth, used for discovery), tasks with a lifecycle (submitted → working → input-required → completed/failed…), messages made of parts, and artifacts as outputs, with streaming and push notifications for long tasks. The A2A docs say explicitly that it complements MCP rather than replacing it, and that it is not a sub-agent or tool-call protocol.
| MCP | A2A | |
|---|---|---|
| Connects | Agent/app ↔ tools, data, prompts | Agent ↔ agent (opaque peers) |
| Counterpart is | A capability provider (stateless-ish function/resource) | An autonomous agent with its own reasoning, possibly long-running |
| Unit of work | Tool call → result | Task with lifecycle, messages, artifacts |
| Discovery | Client lists server's tools | Agent Card |
| Typical use | Give your agent GitHub/Slack/DB access | Your procurement agent delegates to a supplier's quoting agent |
| Main risks | Tool poisoning / prompt injection via tool descriptions and outputs; over-broad credentials | Trust/auth across orgs; untrusted agent outputs; data leakage |
Protocol governance and features moved quickly in 2025–2026: MCP spec revisions (auth, async tasks, elicitation), A2A versions, and foundation membership. Some 2026 reports say A2A has also moved under the AAIF umbrella; I couldn't confirm that from a primary source, so check a2a-protocol.org and the AAIF site for current status.
"MCP vs A2A?" One line: MCP is how an agent uses tools; A2A is how an agent talks to another agent it doesn't control. Then go into security. MCP tool descriptions and outputs are untrusted input (prompt injection, "tool poisoning"), so you need allowlisted servers, least-privilege credentials, and human confirmation for destructive tools. For A2A, cross-org auth and treating peer outputs as untrusted data.
- modelcontextprotocol.io: spec, SDKs, server registry.
- a2a-protocol.org: A2A spec, Agent Cards, task lifecycle.
- Anthropic: Donating MCP and establishing the AAIF: governance context.
16. Failure modes and mitigations
Why agents are hard, quantitatively: if each step succeeds independently with probability \(p\), an \(n\)-step task succeeds with probability \(p^n\). At \(p = 0.95\), 10 steps gives about 60% and 30 steps about 21%; at \(p = 0.99\), 30 steps gives about 74%. Real errors aren't independent and agents can recover, but the intuition holds: long horizons punish small per-step error rates, so verification and recovery matter more than any one step's accuracy.
$$ P(\text{success}) \approx p^{\,n} \qquad 0.95^{10}\approx 0.60,\; 0.95^{30}\approx 0.21,\; 0.99^{30}\approx 0.74 $$| Failure mode | Symptom | Root causes | Mitigations |
|---|---|---|---|
| Infinite loops / repetition | Same tool call over and over; oscillating between two states | Unhelpful error messages; no progress signal; ambiguous done-criteria | Max steps; repeated-call detection; actionable tool errors; explicit finish tool; nudge injection |
| Tool-call errors | Invalid args, wrong tool, hallucinated tool names | Poor tool descriptions; overlapping tools; too many tools | Schema validation + error-as-observation; fewer, distinct tools; examples in descriptions; strict/structured tool modes; tool search for large catalogs |
| Compounding errors | Early wrong assumption poisons everything downstream | No verification between steps; \(p^n\) | Verification gates; checkpoints and rollback; tests; critic with clean context |
| Goal drift | Agent solves a different/easier problem; scope creep | Long contexts dilute the original instruction; distracting observations | Re-state objective in a persistent todo/plan; periodic "am I on task?" checks; restate task in subagent specs |
| Context overflow / context rot | Forgetting earlier facts; degraded reasoning late in long runs | Huge tool outputs; long histories; quality drops with length | Truncate/paginate outputs; compaction; external notes; subagents for exploration (see Memory) |
| Premature completion | "Done!" when it isn't; tests edited to pass | Model's done-judgment ≠ reality | Environment-verified done (tests, checks); forbid editing tests; reviewer agent |
| Over-delegation | Spawning 10 subagents for a simple question; telephone-game loss | Orchestrator not taught to scale effort | Effort-scaling rules in prompt; cap fan-out; single-agent fast path |
| Inter-agent misalignment | Duplicated work; incompatible outputs; agents ignoring each other | Vague task specs; missing shared decisions | Structured specs with decisions_already_made; single writer; shared artifacts |
| Cost blowups | One request costs 100× median | Quadratic history; loops; fan-out; no caching | Budgets per run/tenant; prompt caching; cheaper models for workers; alerting on cost p99 |
| Prompt injection / unsafe actions | Agent follows instructions in a web page or tool output; exfiltrates data | Untrusted content in context + powerful tools | Least privilege; sandbox; HITL on sensitive actions; separate "reader" from "actor"; egress controls |
| Non-reproducibility | Can't debug a failure that happened once | Stochastic sampling; changing tools/data | Full traces (prompts, tool I/O, versions); record/replay; evals on traces |
"How do you know your agent works?" Strong answer: (1) offline evals of end-to-end task success on a representative set, with both final-outcome and trajectory metrics (steps, tool errors, cost); (2) environment-based graders where possible (tests, DB state) and LLM judges with rubrics elsewhere; (3) tracing every run in production; (4) online metrics such as resolution rate, escalation rate, cost/latency percentiles; (5) regression evals on every prompt/model/tool change. Mention τ-bench-style evaluation of tool–agent–user interaction Yao+ 2024 and measuring consistency across repeated trials (pass^k), not just pass@1.
17. How to choose an architecture
Work through these questions in order. Each "no" pushes you toward a simpler design.
| Scenario | Recommended architecture | Why | Watch out for |
|---|---|---|---|
| Extract fields from invoices | Single call + schema validation (+ retry) | Fixed task; no tools needed | Don't build an agent |
| Customer support (high volume) | Router → per-intent chains; narrow tool agent for investigations; handoffs to specialists; HITL on refunds | Predictable majority, long tail needs tools | Policy compliance; cost per ticket |
| Blog post generation | Pipeline: outline → draft → evaluator–optimizer for style/facts | Known stages; clear criteria | Evaluator echo chamber |
| Deep research report | Orchestrator + parallel search subagents + citation verifier | Breadth-first, read-heavy, exceeds one context | Token cost (≈15× chat); source quality |
| Bug fix in a repo | Single-threaded coding agent; read-only explorer subagents; reviewer with clean context; tests as verifier | Writes must be coherent; reads parallelize | Test tampering; sandbox security |
| Large codebase migration (1000 files) | Plan → per-file/per-module workers in isolated worktrees (mechanical, independent changes) → tests per unit → single integrator | Truly independent writes when changes are mechanical | Shared interfaces: decide them up front and pass them in specs |
| Voice receptionist | Streaming single agent (or handoffs), few fast tools, filler speech | Latency budget ≲ 1 s | Supervisor hops; slow tools |
| Multi-day back-office process with approvals | Durable workflow (Temporal) with LLM steps as activities; HITL signals; sagas | Crash-safety, long waits, side effects | Non-idempotent tools; versioning in-flight runs |
| Cross-company procurement | Your agent + A2A to external agents; MCP for internal tools | Opaque peers need a protocol | Trust, auth, untrusted outputs |
The default ladder is: single call → workflow → single agent → single agent + subagents for reading → (rarely) true multi-agent writing. Climb a rung only when an eval shows the lower rung failing on real inputs, and record the cost multiplier you accepted.
- Anthropic: Building effective agents: the "simplest thing that works" principle.
- Google Research blog: When and why agent systems work: accessible summary of the scaling-agents study.
- OpenAI: A practical guide to building agents: single-agent first, then manager vs decentralized patterns.
Interview question bank
1. What's the difference between a workflow and an agent, and how do you decide which to build?
A workflow orchestrates LLM calls and tools through predefined code paths. An agent lets the LLM decide its own sequence of tool calls and when to stop. The decision comes down to whether you can enumerate the steps for your real input distribution. If you can, a workflow is cheaper, faster, more predictable and easier to test, and patterns like chaining, routing, parallelization and evaluator–optimizer cover a lot of ground. Choose an agent when the number and kind of steps depend on what the agent discovers (debugging, research, investigations), when the environment gives feedback the model can use, and when the task's value justifies higher variance and cost. In practice most production systems are hybrids: a workflow skeleton with agentic nodes for the open-ended parts. I'd decide with evals: build the simplest version, measure its failures, and add autonomy only where it fixes them.
2. Walk me through the anatomy of an agent loop. Where do production concerns live?
The harness builds context (system prompt, tool schemas, history), calls the model, and gets back either a final answer or tool calls. It validates and permission-checks each call, executes it (in parallel if independent) with timeouts, converts results or errors into observations (truncated), appends them, checkpoints state, and loops. Production concerns live in the harness, not the model: budgets (steps, tokens, dollars, wall-clock), loop detection, policy/HITL gates for risky tools, retries for transient failures, error-as-observation for semantic failures, context compaction, checkpointing for resumability, and tracing of every prompt and tool I/O. The stop condition should be explicit, either a finish tool or no tool calls, and ideally backed by environment verification rather than the model's own claim.
3. Compare ReAct, plan-and-execute, ReWOO and LLMCompiler.
They differ in when reasoning happens relative to observations. ReAct interleaves think → act → observe every step: very adaptive, but one LLM call per action and quadratic context growth. Plan-and-execute writes a plan first, executes steps (often with a cheaper executor), and replans on failure, which saves strong-model calls and gives a reviewable plan but risks stale plans. ReWOO plans everything up front with placeholders for tool outputs and never re-reads observations during planning; it reported about 5× token efficiency on HotpotQA but can't adapt to surprises. LLMCompiler plans a DAG of tool calls and executes independent ones in parallel, reporting up to 3.7× lower latency and 6.7× lower cost than ReAct, with replanning support. In practice I'd use a hybrid: coarse plan, ReAct execution, parallel tool calls for independent steps, replan on surprise.
4. Back-of-envelope: an agent averages 25 steps, a 6k-token base prompt and adds 1.5k tokens per step. Roughly how many input tokens per task, and how would you cut it?
Total input ≈ \(nP + \Delta n(n-1)/2 = 25 \times 6\text{k} + 1.5\text{k} \times 300 = 150\text{k} + 450\text{k} \approx 600\text{k}\) tokens per task. At 10k tasks/day that's about 6B input tokens/day, so it matters. Levers: (1) prompt caching on the stable prefix and growing history, which typically discounts cached reads heavily; (2) shrink Δ by truncating, paginating or summarizing tool outputs, which usually dominate; (3) compaction, i.e. summarize older turns once past a threshold; (4) fewer steps through better tools (one tool that does the common 3-step sequence), parallel tool calls, or CodeAct-style multi-op steps; (5) offload exploration to subagents so bulky observations never enter the main thread; (6) route easy tasks to a cheaper model or a workflow.
5. Explain the Cognition vs Anthropic debate on multi-agent systems. Who's right?
Cognition (June 2025) argued against multi-agent systems: subagents don't share the parent's full context, and "actions carry implicit decisions", so parallel workers make conflicting assumptions (their Flappy Bird example). They recommended a single-threaded agent with history compression. Anthropic (a day later) showed an orchestrator with parallel subagents beating a single agent by 90.2% on an internal research eval, explained mostly by spending more tokens across parallel context windows, at about 15× chat token cost. Both are right for different task structures. Research is breadth-first and read-only, so it parallelizes. Coding produces one coherent artifact, so parallel writers conflict. The synthesis, which Cognition's April 2026 follow-up also lands on: keep writes single-threaded and use extra agents for intelligence (exploration, review, a "smart friend" model) rather than parallel actions. Controlled studies (e.g. Kim et al. 2025) also find that multi-agent helps on decomposable tasks and hurts on sequential ones.
6. Design a deep-research agent. What are the components and key decisions?
Orchestrator–worker. A lead agent clarifies scope, writes a research plan to memory (so it survives context limits), and decides effort based on query complexity: one agent for a factoid, several parallel subagents for a broad question. Each subagent gets a structured spec (objective, scope, sources to prefer or avoid, output schema, tool budget) and runs its own search/fetch/read loop with parallel tool calls, returning condensed findings with source URLs and noted gaps. Large content goes to an artifact store, passed by reference. The lead evaluates coverage, launches follow-up rounds if needed, synthesizes, and a citation pass verifies each claim against fetched text. Cross-cutting: per-run token budget split across children, dedup of sources, rate-limit semaphores, tracing. Evaluation uses LLM-judge rubrics (factuality, citation accuracy, completeness, source quality, efficiency) plus human review. Expect high token cost, so it's for high-value queries.
7. Supervisor vs handoffs: when would you use each?
Both route work among specialists, but they differ in who owns the conversation. With a supervisor (or agents-as-tools), a central agent calls specialists and gets control back each time, so it can synthesize across them and enforce policy in one place. The costs are an extra LLM hop per turn and a supervisor bottleneck. With handoffs, the active agent transfers control and the specialist talks to the user directly. That's lower latency and gives narrow prompts, but there's no global view, ping-pong is a risk, and it's weak at multi-domain requests that need synthesis. I'd use handoffs for conversational triage (support, voice) where a branch should be fully owned by a specialist, and a supervisor when the final answer combines multiple specialists' outputs or strict central policy is needed. In OpenAI's Agents SDK, handoffs appear to the model as transfer_to_<agent> tools, and history passed to the receiver can be trimmed with an input filter.
8. What do you put in a subagent's task specification, and why?
A subagent knows nothing beyond its prompt, so the spec has to carry the context that would otherwise be lost. It should include: a precise objective and why it matters (lets the subagent make sensible judgment calls); scope (include/exclude, sources); what's already known (to avoid duplicated work); decisions already made (interfaces, style, conventions), which directly addresses conflicting implicit decisions; constraints (tool-call/token budget, deadline, permissions); the output format (a schema, ideally); a done-criterion; and instructions to report uncertainty and gaps. Anthropic reported that vague one-line delegations caused subagents to duplicate work or miss coverage. Ask for condensed outputs or artifact references, not transcripts, to keep the parent's context clean.
9. If each step of an agent is 97% reliable, how reliable is a 20-step task? What does that imply for design?
Assuming independence, \(0.97^{20} \approx 0.54\), roughly a coin flip. At 40 steps it's about 0.30. Errors aren't truly independent and agents can recover, but the lesson holds: long horizons amplify small error rates. Design implications: shorten horizons (better tools that collapse multi-step sequences, workflows for known parts); add verification gates so errors are caught near where they happen (tests, schema checks, critics with clean context); make steps reversible or checkpointed so you can roll back; persist a plan to fight drift; and measure per-step error types in traces to target the dominant failure. Also evaluate consistency across repeated trials, not just single-run success.
10. How would you make a long-running agent survive crashes and multi-day human approvals?
Treat it as a durable workflow. In a Temporal-style design, the orchestration loop is deterministic workflow code, while every LLM call and tool execution is an activity whose result is recorded in event history. After a crash, a worker replays the history and reuses recorded results instead of re-calling the LLM (which would give different outputs). Human approvals arrive as signals, so the workflow can wait days without holding resources. Side-effecting tools take idempotency keys so retries are safe; transient errors get exponential-backoff retries, and semantic errors go back to the model. Multi-system side effects use sagas with registered compensations. Timeouts sit at every level, and workflow/prompt versions are pinned for in-flight runs. With LangGraph, persistent checkpointers plus interrupts give resumability, but you still have to make pre-interrupt side effects idempotent because the node re-runs on resume.
11. What's CodeAct and when would you prefer it to JSON tool calling?
CodeAct makes executable Python the agent's action space. Instead of one JSON tool call per turn, the model writes code that can call several tools, loop, branch, and keep intermediate values in variables, and it sees the printed output or traceback as the observation. The paper reported up to 20% higher success than JSON/text actions across 17 LLMs. Benefits: fewer turns (lower latency and context growth), natural composition, bulky intermediate data stays out of the context window, and it plays to models' coding strength. Downsides: you need a real sandbox (arbitrary code execution is a security boundary), it's harder to apply per-tool permissions and audit individual actions, and failures are messier. I'd prefer it for data-wrangling and multi-step tool composition in a sandbox, and keep typed JSON tools for sensitive, auditable actions like payments.
12. How do you prevent an agent from looping forever or blowing the budget?
Defense in depth in the harness: hard caps on steps, tokens/dollars and wall-clock time; detection of repeated identical tool calls or repeated errors, with a nudge injected or an abort; actionable tool error messages (a common cause of loops is an error the model can't act on); an explicit finish tool and clear done-criteria; observation truncation and compaction to limit quadratic growth; per-tenant and per-run budgets split across subagents; capped fan-out width; and alerting on cost and step-count percentiles. Offline, track the step-count distribution in evals so regressions show up before production. Prompt caching cuts the cost of the loops you do allow.
13. How would you design a coding agent? What matters most?
A single-threaded agent loop over a repo in a sandbox (container/microVM, no prod secrets, restricted egress). The tools are an agent–computer interface: search (grep/glob with concise output), a file viewer with windowing, an edit tool using exact search/replace (fails loudly on non-unique matches, so the agent must read before writing), a shell for tests/linters with output truncation, and git. Behaviors: reproduce the bug with a test first, make minimal edits, run tests, iterate on failures, present a diff. Planning lives in a todo list; long tasks keep progress files and commits. Subagents are used for read-only exploration and for a clean-context reviewer; writes stay in one thread. Guardrails: permission tiers (auto-run tests; approve network/installs/push), and checks against deleting or weakening tests. SWE-agent's main lesson is that ACI design moves results a lot, so invest in the tools.
14. What's the difference between MCP and A2A? Do you need both?
MCP standardizes how an AI application connects to tools, resources and prompts exposed by servers. It's agent-to-capability, JSON-RPC over stdio or HTTP, and solves the M×N integration problem. A2A standardizes communication between autonomous agents that are opaque to each other, possibly across vendors or organizations. It has Agent Cards for discovery, tasks with lifecycles, messages and artifacts, and streaming for long tasks. You need MCP whenever your agent uses external tools in a reusable way. You need A2A only when delegating to agents you don't control, such as a partner's agent. Within one codebase, subagents are usually just function calls or framework constructs, not A2A. Security differs too: MCP's main risk is untrusted tool descriptions/outputs (injection) and over-broad credentials; A2A's is cross-org trust and auth. MCP is now under the Linux Foundation's Agentic AI Foundation, and A2A was donated to the Linux Foundation by Google.
15. Your multi-agent system produces inconsistent outputs: subagents duplicate work and contradict each other. How do you debug and fix it?
First, trace it: look at the actual task specs the orchestrator sent and what each subagent returned. Most problems are specification failures at the boundary. Fixes: richer structured specs (objective, why, scope boundaries to prevent overlap, decisions already made, output schema); have the orchestrator partition the work explicitly (disjoint sources or questions) and track assignments in a ledger; route shared decisions (interfaces, style) through the orchestrator up front; switch from parallel writers to a single writer with parallel readers; add a synthesis/verification step that detects contradictions and resolves them with evidence; and if coupling is inherent, collapse back to a single agent with compaction. The MAST taxonomy is useful here because it separates system-design, inter-agent misalignment and verification failures.
16. When does planning actually help an agent, and how should the plan be represented?
Planning helps on long-horizon, decomposable tasks, on tasks where a human should approve the approach, and on tasks spanning multiple context windows. It adds overhead on short tasks and goes stale quickly in highly reactive environments. Representation should match the need: a structured todo list the agent re-reads and updates each turn (anti-drift, progress visibility), a DAG when you want parallel execution of independent steps, and plan/progress files plus git for multi-session work, since they survive context resets and are editable by humans. Key properties: each item has a verifiable done-criterion, explicit replan triggers (failure, contradiction, periodic review), and completion is verified by the environment rather than self-reported. Reasoning models plan implicitly, so explicit plans matter most for persistence and coordination.
17. How would you add human-in-the-loop to an agent without making it unusable?
Gate by risk, not by default. Classify tools by reversibility and blast radius: auto-allow reads and reversible actions, require approval for irreversible or external ones (payments, emails, deletions, deploys), and deny some outright. Implement approval as an interrupt: checkpoint state, show the reviewer the proposed action, arguments and the agent's rationale, then resume with approve/edit/reject; editing arguments is often more useful than a binary choice. Add policy thresholds (auto-approve refunds below a limit), timeouts with safe defaults, and batch approvals where possible. Make the gated actions idempotent so resumes can't double-execute. Track approval rates. If humans approve 99.9% without changes, the gate is either unnecessary or is being rubber-stamped, and both are problems.
18. Design a voice customer-service agent. How does the architecture differ from a text agent?
Latency dominates: responses need to start within roughly a second, so everything streams. Choose between a cascaded pipeline (VAD/endpointing → streaming STT → LLM → streaming TTS; modular and inspectable) and a speech-to-speech realtime model (lower latency, less control). Turn-taking needs good endpointing, barge-in handling (stop TTS, truncate the assistant message to what was spoken) and backchannels. Keep the agent graph shallow: a single agent or handoffs rather than a supervisor hop per turn. Use a few fast tools, filler speech or async execution for slow ones, and prefetch likely data (caller ID → account lookup). Prompts favor short, speakable sentences. Guardrails and HITL still apply to money actions, often via transfer to a human. Measure latency percentiles per stage, interruption rates and resolution rates.
19. Should you use a framework like LangGraph or roll your own agent loop?
The core loop is small, maybe a hundred lines on a provider SDK, and writing it yourself gives full visibility and lets you use new provider features (parallel tools, caching, thinking) immediately. What's expensive to build is everything around it: persistence/checkpointing, HITL interrupts, streaming, tracing, retries, durable execution. I'd pick a framework when control flow is graph-shaped and you need those features quickly (LangGraph for explicit state graphs, OpenAI Agents SDK for handoff-centric apps, Claude Agent SDK for file/shell-heavy agents). I'd roll my own for latency-critical or unusual control flow, or when a workflow engine like Temporal already provides durability. Either way, I'd require being able to see the exact prompt at every step, keep tools behind MCP or a clean interface so they're portable, and avoid deep framework lock-in in business logic.
20. What failure modes are specific to multi-agent systems compared with single agents?
Beyond the single-agent failures (loops, tool errors, drift), multi-agent systems add: context loss at boundaries (subagents missing information the parent had); conflicting implicit decisions between parallel writers; duplicated or uncovered work from vague task splits; over-delegation (spawning many agents for simple tasks); error amplification when there's no central verification; ping-pong between agents in handoff topologies; termination problems in peer-to-peer chats; supervisor bottlenecks and straggler latency; and cost multiplication (about 15× chat tokens in Anthropic's research system). Debugging is also harder because failures emerge from interactions. Mitigations are structured specs, single-writer designs, centralized verification, effort-scaling rules, capped fan-out, per-run budgets, and end-to-end tracing with agent/span IDs.
21. How would you evaluate an agent architecture change (e.g., moving from single agent to orchestrator + subagents)?
Build an eval set that reflects the real task distribution, including easy queries (where multi-agent may only add cost) and hard breadth-first ones. Measure end-to-end success with environment-based or rubric-based graders, plus trajectory metrics: steps, tool errors, tokens, dollars, latency p50/p95, and consistency across repeated runs (pass^k), because agents are stochastic. Run several trials per task to get confidence intervals. Compare against a strong single-agent baseline (good tools, compaction, same total budget), since much of multi-agent's gain can come from spending more tokens. Slice results by task type to find where it helps and hurts. Then do a staged rollout with online metrics and cost monitoring. If the gain only shows up on a slice, route only that slice to the expensive architecture.
22. What is a blackboard architecture and when is it a good fit for LLM agents?
Agents coordinate through a shared workspace instead of messaging each other. Each reads the current state, contributes when it can, and a controller (or the agents themselves) decides who acts next based on that state. In LLM systems the blackboard is a structured state object (LangGraph state with reducers is a version of this), a shared document, a database, or a repo. It fits incremental, opportunistic problem solving with many specialized contributors and loose coupling, and the shared state is fully inspectable. Downsides: concurrent writes need conflict handling (single writer per key, reducers, optimistic versioning), the board grows and needs curation, and agents can silently overwrite each other's implicit decisions. I'd use it with a typed schema, append-only or reducer-based updates, and a controller that serializes writes to contested fields.