Track B · AI product engineering

Agent Architectures

An "agent" is an LLM running in a loop, choosing tools and acting on what comes back, until it decides it's done or a limit stops it. Almost every AI engineering interview now has an agent question, either conceptual ("ReAct vs plan-and-execute?"), design-shaped ("design a research agent / support agent / coding agent"), or skeptical ("why not just a workflow?"). This page covers the whole design space: the workflow patterns that should come first, the single-agent loop and its reasoning variants, planning, the main multi-agent topologies and the 2025–2026 argument over when they pay off, plus state, durability, protocols, frameworks and failure modes. The goal is that you can pick an architecture for a given problem and defend it with mechanisms and numbers.

TL;DR: the 8–12 things to be able to say out loud

  • Agent = LLM + tools + loop + state. The model picks the next action, the harness runs it, and the result goes back into context. Autonomy is a spectrum, from a single call to a fixed workflow to a router to a full agent loop.
  • Workflows first. Anthropic's taxonomy: prompt chaining, routing, parallelization (sectioning / voting), orchestrator–workers, evaluator–optimizer. Use an agent only when you can't fix the steps in advance and the extra cost and variance are worth it.
  • The loop needs brakes: max steps, token/$ budgets, wall-clock timeouts, no-progress detection and an explicit "finish" tool. Unbounded loops are the classic production incident.
  • ReAct interleaves thought → action → observation. Plan-and-execute / ReWOO plan up front, which saves tokens and latency but is brittle. LLMCompiler plans a DAG and runs it in parallel. Reflexion / evaluator loops add self-critique. ToT / LATS add search. CodeAct uses code as the action space.
  • Planning helps on long-horizon, decomposable tasks, as long as the plan is a living artifact (todo list / DAG / file) that gets revised as observations come in.
  • Multi-agent topologies: orchestrator–workers, supervisor/router, hierarchical, handoffs/swarm, pipeline, debate, blackboard, market/peer-to-peer. Each trades context isolation and parallelism against coordination cost and lost context.
  • The debate: Cognition (2025) said "don't build multi-agents" because actions carry implicit decisions and subagents lack shared context. Anthropic (2025) reported big gains from parallel subagents on breadth-first research at roughly 15× the tokens of chat. Synthesis: parallelize reading, keep writing single-threaded. Cognition's own 2026 follow-up lands in the same place.
  • Handoff context is the hard part: pass a structured task spec (objective, constraints, output format, budget, what's already known), not a vague one-liner.
  • State and control flow: graph state machines (LangGraph-style), checkpoints per step, human-in-the-loop interrupts, resumability. For long-running work, use durable execution (Temporal-style), idempotent tools and sagas.
  • Protocols: MCP = agent ↔ tools/context. A2A = agent ↔ agent across vendors. Both now sit under Linux Foundation governance.
  • Failure modes: loops, tool errors, compounding error (\(p^n\)), goal drift, context overflow, over-delegation, cost blowups. The mitigations are mostly ordinary engineering: budgets, validation, verification, observability, evals.

1. What an agent is

Here's the working definition most practitioners use: an agent is a system where an LLM dynamically directs its own process and tool usage, keeping control over how it accomplishes the task. A workflow is one where LLMs and tools are orchestrated through predefined code paths. Anthropic 2024 Both are "agentic systems". What separates them is who decides the control flow: your code or the model.

Mechanically, an agent has four parts:

LLM (the policy)

Given the current context (instructions, history, observations), it emits either a final answer or one or more tool calls (structured JSON naming a tool plus arguments). It's a stochastic policy \( \pi_\theta(a_t \mid c_t) \) over actions.

Tools (the action space)

Functions the harness can execute: search, code execution, file edits, APIs, DB queries, other agents. Tool descriptions are part of the prompt, and their quality largely determines how reliable the agent is.

Loop (the harness)

Code that calls the model, executes the requested tools, appends results, and repeats until a stop condition. The model never "runs" anything itself. The harness does, so the harness is where you put permissions, budgets and retries.

State (memory)

The context window (short-term), plus any external state: scratchpads, todo files, vector stores, checkpoints. See Memory. State is what lets the agent resume, hand off, or survive context limits.

The basic building block is the augmented LLM: one model call that can use retrieval, tools and memory. Anthropic 2024 Every pattern below is a composition of augmented-LLM calls with different control flow around them.

The autonomy spectrum

"Agent or not" is the wrong binary. Think of a dial for how much control flow the model owns:

LevelWho decides control flowExamplePredictabilityTypical cost / latency
0. Single callNobody (one shot)Summarize this doc; extract fieldsHighest1 call
1. Chain / workflowYour code, fixed stepsOutline → draft → translateHighk calls, fixed
2. RouterLLM picks a branch; code defines branchesClassify ticket → billing / tech / refund flowHigh1 + branch
3. Tool-calling loopLLM picks tools and when to stop, within one taskSupport agent with lookup/refund toolsMediumvariable, often 3–20 calls
4. Planner + executor / multi-agentLLM decomposes the task, spawns sub-workDeep research, large refactorsLowtens to hundreds of calls
5. Long-running autonomousLLM sets sub-goals across sessions; humans check inBackground coding agents working for hoursLowesthours, many context windows
Intuition

Each step up the dial buys flexibility on open-ended inputs and pays for it in variance, cost, latency and debuggability. The engineering question is always: what's the lowest autonomy level that handles the input distribution I actually have? Most production "agents" are levels 1–3 wearing a level-4 marketing label.

Interview angle

"What is an agent?" is a warm-up, and they're checking whether you separate the model from the harness. A strong answer: an LLM in a loop that emits tool calls; a harness that executes them, enforces permissions and budgets, and feeds back observations; state that persists across steps; and a stop condition. Add the workflow vs agent distinction (who owns control flow) and say you'd default to the simplest level that works.

Go deeper

2. Workflows vs agents: the five workflow patterns

Anthropic's "Building effective agents" post (Dec 2024) became the shared vocabulary for this. It lists five workflow patterns and then the autonomous agent as the sixth, most flexible option. Anthropic 2024 Interviewers expect you to know the names and when each one fits.

1. Prompt chaining In LLM 1 gate LLM 2 Out 2. Routing In Router billing flow tech flow small model (easy) 3. Parallelization (sectioning / voting) In LLM a LLM b LLM c aggregate(code / vote) 4. Orchestrator–workers Orchestrator worker 1 worker 2 worker n? Synthesizer subtasks decided at runtime 5. Evaluator–optimizer In Generator Evaluator Out draft feedback pass 6. Agent (for contrast) LLM Environment action observation until stop
Figure 1. The five workflow patterns plus the autonomous agent, adapted from Anthropic's taxonomy. Highlighted boxes are LLM calls; dashed arrows are optional or runtime-determined paths.

The patterns, mechanically

Prompt chaining. Break the task into fixed sequential steps; each LLM call consumes the previous output. Add programmatic gates between steps (schema validation, length check, a classifier) so a bad intermediate result fails fast instead of contaminating later steps. You spend latency to get accuracy: each call does a simpler job. Use it when the decomposition is known and stable (generate marketing copy → translate it; write an outline → check it against criteria → write the doc).

Routing. Classify the input, then dispatch to a specialized prompt, toolset or model. This gives separation of concerns: optimizing the refund prompt can't degrade the tech-support prompt. It also handles cost routing (easy queries to a small model, hard ones to a frontier model). The router can be an LLM, a fine-tuned classifier or an embedding nearest-neighbor. Routing errors are silent, so log the route and evaluate the router on its own.

Parallelization. There are two flavors. Sectioning splits a task into independent subtasks that run concurrently (one call answers the user while another screens for policy violations; one call per document section). Voting runs the same task several times and aggregates. That's self-consistency applied at the system level Wang+ 2022, useful for reviews ("flag if any of 3 reviewers flags") or high-stakes classification. With sectioning, latency is the slowest branch instead of the sum. With voting, cost scales linearly with the number of votes.

Orchestrator–workers. A central LLM decomposes the task at runtime, delegates subtasks to workers, and synthesizes the results. The difference from parallelization is that the subtasks aren't predefined; they depend on the input (which files need changing, which sub-questions need researching). This is already a small multi-agent system, and §6 develops it further.

Evaluator–optimizer. One call generates, another evaluates against explicit criteria and returns feedback, and the loop repeats until it passes or hits a max-iteration cap. It works when (a) clear evaluation criteria exist and (b) feedback measurably improves the output, the way a human editor's would. Literary translation, code that must pass tests, and search tasks that need several rounds are good fits. It's the workflow version of Self-Refine Madaan+ 2023 and Reflexion (§4).

PatternUse whenAvoid whenCost / latency shape
Prompt chainingFixed, known decomposition; each step checkableSteps depend heavily on the contentSum of k calls; sequential latency
RoutingDistinct input categories needing different handling; cost tiersCategories overlap heavily; misroutes are expensive and invisible+1 cheap call; can reduce total cost
Parallel: sectioningIndependent subtasks; guardrails alongside main taskSubtasks share hidden dependenciesCost = sum; latency = max
Parallel: votingNeed confidence or recall; cheap-ish callsCorrelated errors (all samples wrong the same way)n× cost; latency ≈ 1 call
Orchestrator–workersSubtasks unknown until runtime (multi-file edits, research)The decomposition is actually fixed (use chaining)Variable; 1 plan + n workers + synth
Evaluator–optimizerClear criteria, iterative improvement measurableEvaluator no better than generator (self-grading echo chamber)2× per iteration × iterations
AgentOpen-ended, unpredictable number of steps, trusted environment, feedback availableLow tolerance for variance; steps knowableUnbounded unless budgeted
Common mistake

Reaching for a multi-agent framework when a 3-step chain with validation gates would do. Interviewers notice. The canonical advice is to start with direct API calls and the simplest composition, and add complexity only when it demonstrably improves outcomes on your evals. Frameworks "often create extra layers of abstraction that can obscure the underlying prompts and responses". Anthropic 2024

Interview angle

A typical probe: "Design a system to answer customer emails." The weak answer goes straight to "a multi-agent system with a planner…". A strong answer starts with routing (intent classification → per-intent chain), adds parallel sectioning for a policy/PII guardrail, uses evaluator–optimizer for tone/accuracy on drafts, and keeps a narrow tool-calling agent only for the long-tail "investigate account" intent. Then it says how you'd measure whether the agent branch earns its cost.

Go deeper

3. Anatomy of the agent loop

Every agent framework, however much it abstracts, implements something like the loop below. You should be able to write it on a whiteboard and point out where each production concern lives.

def run_agent(task, tools, llm, limits):
    messages = [system_prompt(tools), user(task)]
    state = State(steps=0, tokens=0, cost=0.0, started=now(), seen_actions=Counter())

    while True:
        # ---- budget / stop checks (harness-enforced, NOT model-enforced) ----
        if state.steps >= limits.max_steps:         return finalize(messages, reason="max_steps")
        if state.cost  >= limits.max_cost_usd:      return finalize(messages, reason="budget")
        if now() - state.started > limits.timeout:  return finalize(messages, reason="timeout")
        if context_tokens(messages) > limits.ctx_soft_cap:
            messages = compact(messages)            # summarize / drop old tool outputs

        # ---- THINK: model proposes next action(s) ----
        resp = llm(messages, tools=tools)           # may contain text + 0..n tool calls
        state.update_usage(resp.usage)
        messages.append(assistant(resp))

        if not resp.tool_calls:                     # model chose to answer
            return resp.text                        # (or require an explicit finish() tool)

        # ---- ACT: execute tool calls (parallel if independent) ----
        results = []
        for call in resp.tool_calls:                # or asyncio.gather for parallel calls
            if not policy.allows(call):             # permissions / HITL approval
                results.append(tool_error(call, "denied by policy")); continue
            state.seen_actions[hash(call)] += 1
            if state.seen_actions[hash(call)] > 2:  # loop detection
                results.append(tool_error(call, "repeated identical call; try something else")); continue
            try:
                out = tools[call.name].run(validate(call.args), timeout=limits.tool_timeout)
                results.append(tool_result(call, truncate(out, limits.max_obs_tokens)))
            except Exception as e:                  # errors go BACK to the model as observations
                results.append(tool_error(call, str(e)))

        # ---- OBSERVE: feed results back; loop ----
        messages.extend(results)
        state.steps += 1
        checkpoint(state, messages)                 # durability / resumability
Task+ system prompt ThinkLLM reasons over context Actemit tool call(s) Environmentharness runs tool Observeresult appended to context done?or limit no tool call → answer final answer Each iteration: 1 LLM call (input = entire history so far) + 0..n tool executions
Figure 2. The observe → think → act loop (ReAct shape). The harness, not the model, owns execution, budgets and the stop decision.

Stop conditions and budgets

An agent ends for one of two kinds of reasons: it says it's done (natural), or the harness stops it (forced). You need both.

Stop conditionMechanismNotes
Natural finishModel returns text with no tool call, or calls an explicit finish(answer) / submit toolAn explicit finish tool with a schema makes completion detectable and the output structured. Models sometimes "finish" prematurely, so pair it with verification.
Max steps / turnsCounter in harnessTypical caps range from about 10 (support) to hundreds (coding). Pick from the eval-trace distribution, e.g. p99 steps of successful runs × 1.5.
Token / $ budgetSum usage per callCost grows superlinearly with steps because each call resends the history (see below). Prompt caching helps a lot.
Wall-clock timeoutDeadlineEssential for user-facing latency SLOs and for detecting hung tools.
No-progress / loop detectionHash recent (tool, args); detect repeated errors; compare state snapshotsRespond by injecting a nudge ("you've tried X 3 times"), escalating, or aborting.
Verification gateTests pass, schema valid, judge approvesTurns "I think I'm done" into "the environment confirms done".
Human interruptApproval required for risky actions; user cancelsSee §10 on HITL.

The quadratic cost of loops

Each step resends the whole history. If every step adds about \(\Delta\) tokens (thought + tool call + observation) on top of a base prompt \(P\), then step \(t\) has input \(P + t\Delta\), and total input tokens over \(n\) steps are

$$ \text{Input}_{\text{total}} \approx nP + \Delta \frac{n(n-1)}{2} = O(nP + n^2\Delta). $$

Back-of-envelope: \(P = 5\text{k}\), \(\Delta = 2\text{k}\), \(n = 30\) gives \(150\text{k} + 870\text{k} \approx 1\text{M}\) input tokens for one task. That's why (a) prompt caching matters so much for agents, since the stable prefix is re-read at a large discount; (b) truncating or summarizing tool observations pays off (\(\Delta\) is usually dominated by tool output); and (c) subagents can be cheaper per useful token than one long thread, because their large intermediate observations never enter the parent's context.

Common mistake

Throwing on tool errors. A failed tool call should usually go back to the model as an observation (with an actionable message: "file not found: use an absolute path; did you mean /src/app.py?") so it can self-correct. Only harness-level faults (budget, policy, crash) should end the loop. Equally bad: returning 50k tokens of raw HTML or logs as an observation. Truncate, paginate, or summarize.

Interview angle

"Your agent sometimes runs forever / costs $40 on one request. What do you do?" Strong answers stack several layers: hard caps (steps, $, time) in the harness; loop detection on repeated identical calls; observation truncation and context compaction; prompt caching; an explicit finish tool plus verification; per-tenant budgets and alerts; and evals that track the step-count distribution so regressions show up before production.

Go deeper

4. Single-agent reasoning patterns

These patterns differ in when the model reasons relative to acting, how much it plans ahead, and whether it explores alternatives. Most came from 2022–2024 papers. Today's models have absorbed many of them natively (reasoning models "think" between tool calls; frontier models emit parallel tool calls), so treat them as design vocabulary rather than recipes you have to implement by hand.

ReAct: interleaved reasoning and acting

ReAct prompts the model to alternate Thought → Action → Observation steps, so reasoning traces guide which tool to call next and observations ground the following reasoning. Yao+ 2022 The paper showed that combining the two beat reasoning-only chain-of-thought (which hallucinates facts) and acting-only (which can't plan or recover) on QA/fact-checking and interactive tasks. ReAct is the default shape of essentially every tool-calling loop today. Native function calling replaced the text "Action:" parsing, and reasoning models put the "Thought" in hidden thinking tokens.

Plan-and-execute / Plan-and-Solve

First generate a full plan (a list of steps), then execute the steps, often with a cheaper executor model or plain code, and optionally replan after each step or on failure. Plan-and-Solve prompting showed that asking the model to "first devise a plan, then carry it out" reduces missing-step errors compared with zero-shot CoT. Wang+ 2023 In agent systems the pattern is: a planner (strong model) emits steps; an executor (ReAct sub-loop, or a smaller model) runs each one; a replanner looks at results and revises the remaining steps.

ReWOO: reasoning without observation

ReWOO goes further. The planner writes the whole plan with variable placeholders for tool outputs (e.g. #E1 = Search[x]; #E2 = LLM[summarize #E1]), workers fill in the evidence, and a solver composes the answer. Because the planner never re-reads intermediate observations, it avoids resending the growing history. The paper reports about 5× token efficiency and a 4% accuracy gain on HotpotQA. Xu+ 2023 The cost is zero adaptivity: if #E1 comes back empty, the plan doesn't react.

LLMCompiler: plans as parallel DAGs

LLMCompiler borrows from classical compilers. A Function Calling Planner emits a DAG of tool calls with dependencies, a Task Fetching Unit dispatches calls as soon as their inputs are ready, and an Executor runs independent calls in parallel. Against ReAct it reports up to 3.7× latency speedup, 6.7× cost savings and ~9% accuracy improvement. Kim+ 2023 It also supports replanning when results require it. Modern APIs' parallel tool calls (several tool calls in one model turn) are a lightweight, model-native version of the same idea, limited to one "layer" of the DAG per turn.

Question: "Compare the market caps of Apple and Microsoft and compute the ratio" ReAct (sequential, 3 LLM turns + final): LLMCompiler (DAG, 1 plan + parallel exec + join): think → search(AAPL cap) → observe $1 = search("AAPL market cap") ─┐ think → search(MSFT cap) → observe $2 = search("MSFT market cap") ─┼─► $3 = math($1 / $2) ─► join think → math(a/b) → observe ($1 and $2 run concurrently) ┘ answer latency ≈ 4 LLM + 3 tool (serial) latency ≈ 1 plan + max(tool1, tool2) + tool3 + 1 join

Reflexion and self-critique

Reflexion lets an agent learn across trials without weight updates. After a failed attempt, it writes a verbal reflection ("I failed because I didn't check edge case X"), stores it in an episodic memory buffer, and conditions the next attempt on it. It reported 91% pass@1 on HumanEval against 80% for the GPT-4 baseline at the time. Shinn+ 2023 Self-Refine is the single-trial version: generate → self-feedback → refine. Madaan+ 2023

Common mistake

Assuming self-critique always helps. It works best when the critique is grounded in an external signal (test failures, compiler errors, a retrieval check, a stronger or differently-prompted judge). Pure introspection by the same model on the same context often just rubber-stamps or randomly changes correct answers. In design answers, say where the feedback signal comes from.

Tree search: Tree of Thoughts and LATS

Tree of Thoughts (ToT) generalizes CoT into search over "thoughts" (coherent intermediate steps). The model proposes several candidates per step, evaluates them (value prompt or vote), and runs BFS/DFS with backtracking. On Game of 24, GPT-4 with CoT solved 4% of tasks versus 74% with ToT. Yao+ 2023 LATS (Language Agent Tree Search) brings Monte Carlo Tree Search to acting agents: LM-generated value estimates, self-reflections, and real environment feedback guide expansion. It reported 92.7% pass@1 on HumanEval with GPT-4. Zhou+ 2023

The catch is cost: search multiplies LLM calls by the branching factor × depth, often 10–100× a single ReAct run. It also needs either reversible environments (you can't "backtrack" a sent email) or simulation. In 2026 practice, explicit tree search is rare in products. Its role has largely moved into reasoning models' internal thinking and into best-of-n with a verifier (e.g. run n coding attempts in parallel sandboxes, pick the one that passes tests).

CodeAct: code as the action space

Instead of one JSON tool call per turn, the agent writes executable Python that can call multiple tools, loop, branch, and store intermediate values in variables. The interpreter's output (including tracebacks) is the observation. Across 17 LLMs, CodeAct outperformed JSON/text action formats with up to 20% higher success rate. Wang+ 2024 Hugging Face's smolagents makes CodeAgent its first-class agent type, with sandboxed execution via E2B, Modal, Docker and others. HF smolagents docs

# One CodeAct step replaces ~N JSON tool-call turns:
results = [search(f"{city} population 2025") for city in ["Tokyo", "Delhi", "Shanghai"]]
pops = {c: parse_number(r) for c, r in zip(["Tokyo", "Delhi", "Shanghai"], results)}
print(max(pops, key=pops.get), pops)      # only this printed summary enters context

Why it works: code is composable (loops and function nesting), models have seen huge amounts of it in pretraining, and intermediate data stays in interpreter variables instead of flowing through the context window. Costs: you need a real sandbox (arbitrary code execution is a security boundary), and failures are harder to attribute than with typed tool calls. The same idea appears as "code execution with MCP" or programmatic tool calling in vendor APIs, where the model writes code that calls tools so large intermediate results never enter context.

PatternReason ↔ act timingLLM callsAdaptivityBest forMain risk
ReActInterleaved, every step1 per actionHighDefault; unpredictable environmentsWandering, loops, quadratic context
Plan-and-executePlan first, replan on eventsFew big + many smallMediumLong multi-step tasks; human plan approvalStale plans
ReWOOAll reasoning up front~2 (plan + solve)LowPredictable tool chains; token savingsCan't react to surprises
LLMCompilerDAG plan, parallel execPlan + join (+ replan)MediumMany independent lookups; latency-sensitiveWrong dependency graph
Reflexion / evaluator loopAfter an attempt× trialsAcross trialsTasks with a checkable outcome (tests)Ungrounded self-critique
ToT / LATSSearch over branches× branching × depthHigh (backtracking)Puzzles, code with verifiers, offlineCost; irreversible actions
CodeActInterleaved, but multi-op per stepFewer turnsHighData wrangling, multi-tool compositionSandbox security; debuggability
Interview angle

"ReAct vs plan-and-execute?" A strong answer frames it as a latency/cost vs adaptivity trade-off on a known axis. ReAct re-decides after every observation (adaptive, expensive). Plan-first saves calls and gives an approvable artifact but needs a replanning trigger. The realistic answer is a hybrid: plan coarsely, execute with a ReAct sub-loop, replan on failure or surprise, and parallelize independent steps (LLMCompiler-style). Bonus: mention that reasoning models blur the line because they plan internally between tool calls.

Go deeper

5. Planning in depth

"Planning" in agents covers three separate decisions: when to plan (upfront vs interleaved), how to represent the plan, and how deep to decompose (flat vs hierarchical).

Upfront vs interleaved replanning

Upfront plan

Produce the whole plan before acting. It's cheap, inspectable and parallelizable (a DAG), and humans can approve it. It works when the environment is predictable and early steps don't change what later steps should be. It fails when reality diverges: a missing file, an API returning nothing, a wrong assumption about the codebase.

Interleaved / replanning

Plan a bit, act, observe, revise. Triggers for replanning include a step failing, an observation contradicting a plan assumption, a periodic checkpoint (every k steps), or new user input. It costs more, but it's robust. Most effective coding and research agents keep a high-level plan and revise it continually.

Plan representations

RepresentationWhat it looks likeProsConsSeen in
Free-text plan in context"1. find config 2. patch 3. test"Zero infraGets lost as context grows; not machine-checkablePlan-and-Solve prompts
Structured todo list (tool-managed)todo_write([{id, content, status}])Model re-reads and updates it; progress visible to UI and user; fights goal driftStill flat; model can mark items done falselyCoding agents (e.g. Claude Code's todo tool)
DAG of tasksNodes = tool calls/subtasks, edges = data depsParallel execution; explicit dependenciesHarder for the model to produce correctly; rigidLLMCompiler, ReWOO, HuggingGPT
Plan/progress files on diskPLAN.md, progress.txt, features.json + git historySurvives context resets and sessions; humans can edit; other agents can readNeeds discipline to keep it updatedLong-running harnesses Anthropic 2025
Hierarchical task networkGoal → subgoals → primitive actionsScales to big tasks; maps onto multi-agent delegationDecomposition errors propagate down the treeHierarchical multi-agent, HuggingGPT-style task planning Shen+ 2023

Anthropic's long-running-agent harness (Nov 2025) is a good concrete example of "plan as files". An initializer agent runs once to set up the environment: an init.sh, a claude-progress.txt log, an initial git commit, and a JSON feature list where every feature starts with a "passes" status of false. Each later coding agent session reads the progress file and git log, picks one feature, implements it, verifies it end to end (browser automation), commits, and updates progress. Anthropic 2025 The plan lives outside any single context window, so the work survives context resets.

Hierarchical decomposition

For big tasks, decompose recursively until the leaves are "one agent, one context window, verifiable". Rules of thumb:

When does planning help?

Interview angle

"How does your agent plan?" Interviewers want to hear about plan representation and lifecycle, not just "it makes a plan". A strong answer: a structured todo list the agent re-reads every turn, explicit replan triggers, done-criteria per item, and persistence outside the context window for long tasks. Plus awareness that the agent may mark items done that aren't, which is why done is verified by the environment, not by the model's say-so.

Go deeper

6. Multi-agent architectures

A "multi-agent system" is any system where more than one LLM loop, each with its own prompt, tools and (usually) its own context window, cooperates on a task. The reasons to split are always some mix of: context isolation (each agent sees only what it needs), parallelism (wall-clock speed, more total tokens spent on the problem), specialization (different prompts/tools/models per role), and organizational modularity (teams own different agents). The costs are always: lost context at boundaries, coordination overhead, more tokens, and harder debugging.

Intuition

Think of agents as microservices whose API is natural language. All the distributed-systems lessons apply: chatty interfaces are slow, shared mutable state needs coordination, the network (here, the context handoff) is lossy, and "we split it into services" doesn't fix a bad domain model. The extra wrinkle is that each service is also non-deterministic and confidently fills gaps in its spec with guesses.

(a) Orchestrator–worker / lead agent with subagents

A lead agent plans, spawns subagents with specific subtasks (often in parallel), receives their condensed results, and synthesizes. Subagents usually can't talk to each other. Everything goes through the lead.

User query Lead agentplan → save to memoryspawn, evaluate, iteratesynthesize Subagent 1own context · search tools Subagent 2own context · search tools Subagent nown context · search tools Citation agentattach sources Memory (plan, notes) task spec →← condensed findings final draft
Figure 3. Orchestrator–worker as in Anthropic's multi-agent research system: a lead agent fans out to parallel subagents with isolated contexts, then a citation pass. Dashed arrows carry compressed results, not full transcripts.

The reference implementation is Anthropic's Research feature (June 2025). A lead agent (Claude Opus 4) plans, saves the plan to memory (in case the context gets truncated), spawns subagents (Claude Sonnet 4) that search in parallel, then synthesizes, with a separate CitationAgent pass to attribute sources. Anthropic 2025 The reported numbers, worth memorizing with their caveats:

Engineering lessons from that post that come up in interviews: teach the orchestrator to scale effort to query complexity (simple fact-finding gets 1 agent with a few tool calls; complex research gets many subagents); give subagents detailed task descriptions (objective, output format, tool guidance, boundaries), because vague delegation caused duplicated work and gaps; use "start wide, then narrow" search strategies; and have subagents write outputs to a filesystem/artifact store and pass back lightweight references, so large content doesn't get squeezed through the lead's context. Anthropic 2025

Pros

  • Parallel breadth: wall-clock speed and more total tokens on the problem
  • Context isolation: each subagent's window is clean and focused; the lead stays small
  • Natural fit for "read-heavy" tasks (research, search, code exploration, review)
  • The lead keeps one coherent decision-making thread

Cons

  • About 15× chat tokens; only worth it for high-value tasks
  • The lead is a bottleneck and single point of failure; synchronous waves block on the slowest subagent
  • Subagents lack each other's context, so you get duplicated work or inconsistent assumptions
  • Lossy compression of findings back to the lead

(b) Supervisor / router

A supervisor agent decides, at each turn, which specialist agent acts next (or routes the whole request once). Specialists return control to the supervisor after they act. It differs from orchestrator–worker mostly in emphasis: the supervisor is a dispatcher over a fixed roster of specialists (often sequential), whereas the orchestrator creates subtasks dynamically (often parallel). In OpenAI's Agents SDK terms this is the "agents as tools" / manager pattern: the manager calls specialists as tools and keeps ownership of the final answer. OpenAI docs

(c) Hierarchical teams

Supervisors of supervisors: a top-level coordinator delegates to team leads (e.g. "research team", "writing team"), each managing their own workers. It mirrors org charts and is how MetaGPT-style "software company" simulations are structured, with roles like product manager, architect and engineer following SOPs. Hong+ 2023

(d) Handoffs / swarm

Decentralized transfer of control: the active agent decides to hand the conversation to another agent, which then becomes the active agent and talks to the user directly. There's no central supervisor in the loop. OpenAI's experimental Swarm library introduced this pattern, and its README now says it's replaced by the OpenAI Agents SDK, "a production-ready evolution of Swarm". OpenAI Swarm In the Agents SDK, handoffs are exposed to the model as tools: a handoff to an agent named "Refund Agent" appears as a tool named transfer_to_refund_agent. By default the receiving agent sees the full conversation history, and an input_filter can trim what it sees. OpenAI Agents SDK docs

Supervisor (hub & spoke) User Supervisor Agent A Agent B Agent C control always returns to the hub Handoffs (swarm) User Triage Billing Refunds transfer active agent talks to user; control moves sideways Hierarchical teams Top coordinator Research lead Writing lead w1 w2 w3 w4 intent compressed at every level
Figure 4. Supervisor vs handoff vs hierarchical. The question each answers: who owns the conversation, and where does control go after a specialist acts?

(e) Sequential pipelines / assembly line

Fixed stages, each an agent (or LLM call) with its own role: researcher → writer → editor → fact-checker, or localize → repair → validate for code. This is really prompt chaining where the stages happen to be agentic. Agentless is a well-known example: a fixed three-phase localization → repair → patch-validation process with no LLM-chosen control flow. It scored 32% on SWE-bench Lite at about $0.70 per issue, beating the open-source agents of the time. Xia+ 2024

(f) Debate / multi-agent critique

Several agents (same or different models, possibly with assigned stances) propose answers, read each other's reasoning, and revise over rounds until they converge or a judge decides. Du et al. showed that multi-agent debate improved math and strategic reasoning and reduced factual errors. Du+ 2023 The generator–verifier loop is the practical cousin: a reviewer agent with a clean context checks the producer's output.

(g) Blackboard / shared state

A classic AI architecture: agents don't message each other. They read from and write to a shared workspace (the "blackboard"), and a controller (or the agents themselves) decides who acts next based on the blackboard's state. In LLM systems the blackboard is a shared document, a structured state object (LangGraph's graph state is effectively this), a database, or a repo.

(h) Market / auction, peer-to-peer and network topologies

Market/auction (contract-net): a manager announces a task, agents bid (with estimated cost, confidence or capability), and the best bid wins. It's useful when agents are heterogeneous (different costs/tools) and the roster is dynamic, e.g. cross-organization agents discovered via A2A agent cards. Peer-to-peer / network: any agent can message any other (AutoGen-style group chats, CAMEL role-playing Li+ 2023). It's flexible, but communication grows as \(O(n^2)\), it's hard to reason about termination, and it's mostly a research or simulation setup. Generative Agents (the "Smallville" simulation) is the famous example of agent societies used for simulation rather than task completion. Park+ 2023

May be out of date

As of 2026, practitioner consensus (including Cognition's April 2026 follow-up) is that unstructured swarms, meaning arbitrary networks of agents negotiating with each other, mostly don't see meaningful production adoption, while structured manager/verifier patterns do. Cognition 2026 Market-based coordination among LLM agents is still largely research. Check the current literature before presenting it as production practice.

Summary table

TopologyControlCommunicationParallelismBest forBiggest risk
Orchestrator–workersCentral, dynamic decompositionTask spec down, summary upHighBreadth-first research, codebase exploration, reviewSubagents lack shared context; 15× tokens
Supervisor / routerCentral dispatcherHub and spokeLow–mediumFixed set of specialists; policy centralizationExtra hop latency; supervisor bottleneck
HierarchicalTree of supervisorsUp/down the treeMedium–highVery large decomposable tasksIntent loss per level
Handoffs / swarmDecentralized; active agent transfersShared conversation passed alongNone (one active)Conversational triage, customer supportPing-pong; no global view
Sequential pipelineFixed by codeStage outputsNone (but can pipeline across items)Repeatable processes, content pipelinesEarly errors propagate; rigid
Debate / critiqueRounds + judgeBroadcast of proposalsHigh per roundHigh-stakes answers, review, verificationCost; conformity
BlackboardController or opportunisticShared state, no direct messagesMediumIncremental problem solving; many contributorsWrite conflicts; state bloat
Market / P2PEmergent / biddingAny-to-anyHighHeterogeneous cross-org agents; simulationsTermination, \(O(n^2)\) chatter, unpredictability
Interview angle

"Supervisor vs handoff?" Ask who should own the final answer. If one agent must synthesize across specialists and keep consistent policy, use a supervisor or agents-as-tools. If a specialist should take over the conversation for a branch (e.g. refunds) and talk to the user directly, use a handoff. OpenAI's own docs frame the choice exactly this way, and also say "start with one agent whenever you can". OpenAI docs

Go deeper

7. Agent roles and how they differ

Role names get used loosely in interviews and papers. Here's a precise vocabulary. One agent can play several roles, and a role can be plain code instead of an LLM.

RoleResponsibilityOwnsOften implemented asDon't confuse with
PlannerTurns a goal into steps / a DAG; replans on eventsThe plan artifactStrong reasoning model, called rarelyOrchestrator (planner doesn't execute or delegate)
OrchestratorDecomposes at runtime, spawns workers, integrates resultsTask lifecycle + synthesisLead agent with a "spawn subagent" toolRouter (one-shot dispatch, no synthesis)
Coordinator / supervisorChooses which existing agent acts next; enforces turn-taking and policyControl flow among a fixed rosterLLM with "route_to(agent)" tool, or codeOrchestrator (creates new subtasks)
RouterClassifies input once and dispatchesA single routing decisionSmall model / classifier / embeddingsSupervisor (multi-turn control)
Executor / workerPerforms a bounded subtask with toolsIts subtask's outputReAct loop with narrow tools; cheaper modelPlanner
Critic / verifierChecks outputs against criteria; returns pass/fail + feedbackQuality gateJudge LLM with clean context, tests, linters, schema validatorsGenerator self-review (same context, correlated errors)
Memory managerDecides what to persist, summarize, retrieve, forgetLong-term store + compactionBackground LLM job or code policy (see Memory)Blackboard (shared working state, not long-term memory)
Synthesizer / reporterMerges worker outputs into a final artifactFinal answerOften the orchestrator itselfCitation agent (post-hoc attribution)
Common mistake

Making every role a separate LLM agent. A verifier that's really "run the test suite" should be code. A router over 4 well-separated intents can be a fine-tuned classifier at about 1/100th the latency. Use an LLM for a role only when the role needs judgment.

8. The case against vs the case for multi-agent

This was the defining architecture debate of mid-2025, and interviewers love it because it tests judgment rather than recall. The two key posts came out one day apart in June 2025.

Against: Cognition, "Don't Build Multi-Agents" (Walden Yan, 12 Jun 2025)

Two principles Yan 2025:

  1. Share context, and share full agent traces, not just individual messages. A subagent given only a task description lacks the decisions and nuance in the parent's history.
  2. Actions carry implicit decisions, and conflicting decisions carry bad results. Parallel agents make unstated assumptions (style, interfaces, approach) that clash when combined.

Example: "build a Flappy Bird clone" split into "background" and "bird" subtasks. One subagent builds a Super Mario-style background, the other a bird that doesn't match, and the combiner has to reconcile incompatible work. Recommendation: a single-threaded linear agent, and for very long tasks a dedicated model that compresses history into key details and decisions.

For: Anthropic, multi-agent research system (13 Jun 2025)

For breadth-first research queries with many independent directions, parallel subagents with isolated contexts won big (+90.2% on the internal eval). Token spend explains most of the variance, and multi-agent is a way to spend more tokens usefully by spreading them across many context windows in parallel. Anthropic 2025

The same post is explicit about limits: multi-agent fits less well for domains that need all agents to share the same context or that have many dependencies between agents, and it notes that most coding tasks have fewer truly parallelizable parts than research. It's also expensive (about 15× chat tokens), so the task's value has to justify it.

The synthesis

Read the two posts together and they mostly agree. LangChain's commentary at the time framed it as reading vs writing: multiple agents gathering information (reading) parallelize well, while multiple agents producing parts of one artifact (writing) create conflicting implicit decisions that are hard to merge. Chase 2025 Cognition's own April 2026 follow-up, "Multi-Agents: What's Actually Working", lands on the same principle: writes stay single-threaded, and additional agents contribute intelligence rather than actions. The patterns it reports working are a clean-context code-review agent (generator–verifier), a "smart friend" frontier model consulted by the primary agent, and manager agents that coordinate scoped child agents. Parallel-writer swarms still don't work well. Cognition 2026

Controlled research points the same way. A Dec 2025 study by Google researchers and collaborators evaluated 260 configurations across six agentic benchmarks and five architectures (single-agent; independent, centralized, decentralized and hybrid multi-agent). Multi-agent performance relative to single-agent ranged from +80.8% on decomposable financial reasoning to −70.0% on sequential planning. Coordination showed diminishing returns once the single-agent baseline was already strong, tool-heavy tasks paid a multi-agent overhead, and architectures without centralized verification propagated errors more. Kim+ 2025 The MAST failure taxonomy (1,600+ annotated traces across 7 multi-agent frameworks) groups failures into system design issues, inter-agent misalignment, and task verification, and notes that multi-agent gains on popular benchmarks are often minimal. Cemri+ 2025

Multi-agent tends to win when…Single agent tends to win when…
The task is breadth-first / embarrassingly parallel (many independent sub-questions, many files to read)The task is sequential / tightly coupled (each step depends on the last; one coherent artifact)
Subtasks are read-only or produce independent outputsSubtasks write to shared artifacts with implicit design decisions
Total information exceeds one context window, and subagents can compress itEverything fits in one context (perhaps with compaction)
Wall-clock latency matters and parallel spend is affordableCost per task matters more than latency
You want an independent verifier with clean contextCoordination overhead would dominate (simple tool-heavy tasks)
Task value is high (research reports, large migrations)High-volume, low-value requests (support FAQs)
Interview angle

If asked "single or multi-agent?", don't pick a side. A strong answer goes: "It depends on task structure. I'd parallelize the read-heavy, independent parts (research, exploration, review) with isolated subagents returning condensed results, and keep any writes to a shared artifact in a single thread that owns the decisions. Expect about 15× tokens versus chat, so only for high-value tasks, and I'd validate the gain on my own evals against a strong single-agent baseline with compaction." Name both posts and the Cognition 2026 update.

Go deeper

9. Inter-agent communication and handoff context

Agents communicate in three basic ways, and the choice decides what context survives a boundary.

MechanismHowProsConsExample
Shared transcriptAll agents read/append one message historyMaximum shared context; simpleContext bloat; agents distracted by irrelevant turns; prompt confusion about "who am I"Group chat; Agents SDK handoffs (full history by default)
Message passingAgent A sends a structured message/task to B; B returns a resultIsolation; small contexts; clear contractsLossy: B only knows what A wrote downSubagent-as-tool; A2A tasks
Artifacts / files / shared stateAgents write to files, DB, state object; pass referencesLarge content doesn't squeeze through contexts; persistent; inspectableNeeds schemas, versioning, conflict handlingResearch subagents writing to a filesystem; LangGraph state; repos

What to pass in a handoff

Most multi-agent failures are really specification failures at the boundary. Every subagent starts with zero knowledge except its prompt. A good structured task spec looks like this:

{
  "objective": "Find 2024-2026 benchmark results for open-weight models on SWE-bench Verified",
  "why": "Parent is writing a comparison table for a CTO memo; needs numbers + sources",
  "scope": {"include": ["official leaderboards", "model cards"], "exclude": ["blog speculation"]},
  "known_context": ["Parent already has results for closed models; don't repeat them"],
  "constraints": {"max_tool_calls": 15, "max_tokens_out": 1200, "deadline_s": 120},
  "output_format": {"type": "json", "schema": "[{model, score, date, source_url}]"},
  "done_when": ">= 5 models with primary-source URLs, or budget exhausted (report gaps)",
  "decisions_already_made": ["Use 'Verified' split only", "Report pass@1"]
}

Notice why and decisions_already_made. They're the direct mitigation for Cognition's "implicit decisions" problem: you write down the decisions the subagent would otherwise make differently. Also ask for condensed outputs with references (findings + source pointers, or a file path) rather than raw transcripts, so the parent's context stays clean.

Returning results

Common mistake

Handoffs that pass the full conversation to a narrow specialist. The specialist's instructions get diluted by thousands of irrelevant tokens and earlier agents' persona text. Filter history at handoff time: the Agents SDK exposes input_filter for exactly this. OpenAI Agents SDK docs The opposite failure, passing a one-line task, is just as common. The right amount is a structured spec plus relevant excerpts.

10. State management and control flow

Once an agent does more than one tool loop, you need to answer: where does state live, who decides the next step, and what happens if the process dies or a human must approve something? The dominant answer since 2024 is to model the agent as an explicit state graph.

Graph-based state machines (LangGraph-style)

In LangGraph, you define a typed state object, nodes (functions, which may call LLMs or tools, that read state and return updates), and edges (fixed or conditional, where a function of state picks the next node). Updates are merged into state via per-key reducers (e.g. append to the message list rather than overwrite), which is what makes parallel branches safe to join. Execution proceeds in "super-steps", and with a checkpointer attached a state snapshot is saved at every super-step, keyed by a thread_id. LangGraph docs The same model covers fixed workflows (all edges fixed), ReAct agents (an LLM node ↔ tool node cycle with a conditional "tool calls?" edge), and multi-agent systems (nodes are subgraphs).

START planLLM node agentLLM w/ tools toolsexecute calls toolcalls? human_reviewinterrupt() verifytests / judge END yes (safe) risky approved / edited no → done? pass fail → replan State {messages[], plan, artifacts,step_count, approvals} checkpointedafter every super-step
Figure 5. A LangGraph-style state graph: plan → agent ⇄ tools loop, a conditional edge that routes risky actions to a human interrupt, and a verify node that either ends or sends the agent back to replanning.

Checkpoints, human-in-the-loop and resumability

Common mistake

Treating "checkpointed state" as "durable execution". A checkpointer saves state between steps. It doesn't by itself guarantee that a step interrupted halfway through a non-idempotent side effect is handled correctly, or that a crashed process gets rescheduled automatically. For that you need a durable execution engine or your own retry and idempotency discipline (next section).

Interview angle

"How would you add human approval to an agent that can issue refunds?" Strong answer: model the refund as a tool whose execution routes through an approval node; persist state at the interrupt (checkpointer keyed by conversation/thread); notify a reviewer with the proposed arguments plus the agent's rationale; resume with approve/edit/reject; make the refund call idempotent with a key derived from (ticket, amount) so a resume or retry can't double-pay; set a timeout path if no one approves. Bonus: policy thresholds, e.g. auto-approve under a set amount.

Go deeper

11. Durable execution for long-running agents

An agent that runs for minutes to hours, waits on humans for days, or calls flaky external APIs is a long-running distributed workflow, and it should be built like one. The lessons come straight from workflow engines like Temporal.

The Temporal model, applied to agents

Temporal splits code into workflows (deterministic orchestration code whose progress is recorded in an event history and recovered by replay after a crash) and activities (the non-deterministic, side-effecting work: network calls, DB writes, and, for agents, LLM calls and tool executions), which get automatic retries with configurable policies. Temporal docs Temporal docs Mapped onto an agent:

Workflow (deterministic, replayable) Activities (side effects, retried) ───────────────────────────────────── ────────────────────────────────── loop: call_llm(messages) ← recorded result resp = await activity(call_llm, msgs) run_tool(name, args) ← retry policy, timeout if resp.done: return resp send_email(..., idem_key) for call in resp.tool_calls: charge_card(..., idem_key) if risky(call): await signal("approval") ← human approval arrives as a signal (days OK) r = await activity(run_tool, call) msgs += results On crash: a new worker replays history, and recorded LLM/tool results are reused, not re-executed.

Why this matters: an LLM response is non-deterministic, so it must be an activity whose result is recorded. Otherwise replay would get a different answer and diverge. The workflow keeps only control logic. Temporal has published integrations with agent frameworks (for example a LangGraph plugin) that push LLM and tool calls into activities. Temporal blog

Idempotent tools, retries and sagas

Interview angle

"Your agent runs a 2-hour data migration and the pod gets killed at minute 70. What happens?" A weak answer: "it restarts". A strong answer: state is checkpointed per step (or the run is a durable workflow); completed steps' LLM and tool results are recorded and replayed, not re-executed; the in-flight step's side effect is idempotent (keyed) so re-running it is safe; non-reversible steps were gated; and partial failure triggers compensations. Then mention observability: a run ID tying traces, cost and state together.

Go deeper

12. Concurrency and parallel tool calls

There are three levels of parallelism in agent systems:

  1. Parallel tool calls within one turn. Modern model APIs let the model emit several tool calls in one response (e.g. three searches). The harness runs them concurrently (asyncio.gather) and returns all results together. It's cheap to add and cuts latency for independent lookups. In Anthropic's research system, parallel tool calls in subagents were one of the two parallelization levers behind the "up to 90%" time reduction. Anthropic 2025
  2. Parallel branches in a workflow / graph. Fan out to nodes, fan in with reducers (LangGraph), or use a DAG executor (LLMCompiler).
  3. Parallel agents. Subagents in separate contexts, possibly separate processes or sandboxes.
ConcernWhat goes wrongMitigation
Hidden dependenciesModel parallelizes calls that depend on each other (edit file, then run test, concurrently)Mark tools as read-only vs mutating; serialize mutating calls; prompt guidance
Write conflictsTwo agents edit the same file / recordSingle writer; locks; git worktrees per agent plus merge; optimistic concurrency with versions
Rate limitsFan-out of 20 agents × 5 tools hits provider 429sSemaphores / token buckets per provider; backoff; queue
Straggler latencySynchronous wave waits for slowest subagentPer-subagent timeouts; accept partial results; async/streaming orchestration
Cost amplificationParallelism multiplies spend quicklyBudget per run split across children; cap fan-out width
Ordering in contextResults returned out of order confuse attributionMatch results to call IDs (APIs require tool_use_id / call_id)

Back-of-envelope: a research step with 5 independent searches at about 2 s each takes about 10 s serially and about 2 s in parallel. Five subagents that each run 10 sequential steps of about 6 s (LLM + tool) finish in about 60 s wall-clock instead of about 300 s, at roughly the same or higher token cost. Parallelism buys latency, not cost.

13. Specialized agent types

Coding agents

Coding is the most commercially mature agent domain, because the environment gives dense, objective feedback (compilers, linters, tests) and the work is reversible (git). The typical loop: understand the task → explore the repo (search, read files) → plan → edit → run tests/linters → read failures → iterate → present the diff.

May be out of date

Coding-agent capability, SWE-bench scores, and practices (background/async agents, many parallel agents in git worktrees, AGENTS.md-style repo instructions) changed month to month through 2025–2026. Check swebench.com and vendor posts for current numbers before quoting any.

Computer-use and browser agents

The action space is the GUI: screenshots (plus optionally the accessibility tree or DOM) come in as observations, and mouse/keyboard actions go out. Anthropic, OpenAI and Google all ship computer-use or browser-use model capabilities. Claude computer use docs Key engineering points:

Deep-research agents

The canonical orchestrator–worker use case: plan sub-questions → parallel search/read subagents → iterative deepening ("start wide, then narrow") → synthesis → citation verification. Design concerns: source quality ranking (prefer primary sources over SEO content), deduplication across subagents, claim-level citations checked against the fetched text, explicit "what I couldn't find" sections, and scaling effort to the question (don't spawn 10 agents for a factoid). Anthropic 2025 Evaluation is hard because answers are open-ended, so use LLM-judge rubrics (factual accuracy, citation accuracy, completeness, source quality, tool efficiency) plus human spot checks.

Voice agents

Voice agents run under a hard latency budget. In human conversation the gap between turns is short, and responses much slower than about ~1 s start to feel unnatural. That reshapes the architecture:

Agent typeObservationAction spaceFeedback signalDominant constraintTypical architecture
CodingFiles, test output, errorsRead/search/edit/shellTests, compilers (dense, objective)Correctness; safety of shellSingle-threaded writer + read-only explorer/reviewer subagents
Computer-use / browserScreenshots, DOMClick/type/scrollPage state (noisy)Grounding, latency, injectionSingle agent in sandboxed VM; HITL for sensitive steps
Deep researchSearch results, pagesSearch/fetch/readWeak (judge rubrics)Breadth, source quality, costOrchestrator + parallel subagents + citation pass
VoiceTranscribed speech (or audio)Speak + toolsUser reactionLatency (< ~1 s), turn-takingStreaming single agent or handoffs; minimal hops
Customer supportUser messages, CRM dataLookup/refund/escalateResolution, CSATPolicy compliance, cost/volumeRouter + narrow tool agents + HITL on money
Go deeper

14. Frameworks landscape (as of late 2026)

May be out of date

This landscape changes every quarter. Status notes below were checked against official docs in October 2026, but versions, names and consolidations (e.g. AutoGen → Microsoft Agent Framework) keep happening. Check current docs before making claims in an interview.

FrameworkFromCore abstractionMulti-agent styleNotes / 2026 status
LangGraph docsLangChainTyped state graph: nodes, conditional edges, reducersAnything as subgraphs: supervisor, swarm, hierarchicalLow-level and explicit; checkpointers, interrupts, durable execution modes. LangGraph reached a 1.0 release in late 2025. The usual pick when you want explicit control flow.
OpenAI Agents SDK docsOpenAIAgent (instructions + tools), Runner loopHandoffs; agents-as-toolsProduction successor to Swarm; guardrails, sessions, tracing built in. Python and TypeScript.
Claude Agent SDK docsAnthropicClaude Code's agent loop as a librarySubagentsBuilt-in file/shell/web tools, hooks, permissions, sessions, MCP, skills; Python and TypeScript. Formerly the Claude Code SDK.
CrewAI docsCrewAIRole-based "crews" of agents with tasks; "Flows" for event-driven controlSequential or hierarchical (manager) processesHigh-level, role/goal/backstory framing; quick to prototype.
Microsoft Agent Framework docsMicrosoftAgents + graph-based workflowsOrchestration patterns inherited from AutoGen (sequential, concurrent, group chat, handoff, magentic)Successor that unifies AutoGen and Semantic Kernel; .NET and Python. Public preview Oct 2025, 1.0 GA reported Apr 2026. MS blog AutoGen is in maintenance mode.
AG2 GitHubCommunity (AutoGen fork)Conversable agents, group chatConversation-drivenCommunity continuation of the original AutoGen 0.2 line.
Google ADK docsGoogleLlmAgent + workflow agents (Sequential, Parallel, Loop); graph workflows in 2.0Agent hierarchies; A2A integrationPython, TypeScript, Go, Java, Kotlin; tight Gemini/Vertex integration but model-agnostic.
LlamaIndex docsLlamaIndexEvent-driven Workflows; agents over data/RAGMulti-agent workflowsStrongest on document/data-centric agents and retrieval.
Pydantic AI docsPydanticType-safe Agent with typed deps and structured outputsDelegation via tools, programmatic handoff, graphs"FastAPI feel"; strong validation; model-agnostic.
smolagents docsHugging FaceCodeAgent (code as action) and ToolCallingAgentManaged agentsMinimal (core logic around a thousand lines); sandboxed execution options.
DSPy docsStanford NLPDeclarative modules (signatures) + optimizers that compile prompts/few-shotsCompose modules (e.g. dspy.ReAct)Not an orchestration framework so much as a way to optimize LM programs against a metric. Khattab+ 2023

Framework vs roll your own

Use a framework when…

  • You need checkpointing, HITL interrupts, streaming, and tracing now, not after a quarter of infra work
  • The team is new to agents and benefits from opinionated structure
  • You want ecosystem integrations (MCP, vector stores, observability)
  • Your control flow is graph-shaped and you'd otherwise reinvent a state machine

Roll your own when…

  • The core loop is ~100 lines and you want full visibility into every prompt and token
  • You need tight latency (voice) or unusual control flow
  • You already have a workflow engine (Temporal, Step Functions) for durability
  • Framework abstractions fight your model provider's newest features (parallel tools, caching, thinking)

The common middle path: a thin in-house loop on the provider SDK, a real workflow engine for durability, MCP for tool integration, and OpenTelemetry-style tracing. Whichever you pick, make sure you can see the exact prompt sent to the model at every step. That's the debugging superpower frameworks sometimes hide.

Interview angle

"Which framework would you use?" isn't a trivia question. They want trade-off reasoning. A good answer: "For an explicit, auditable workflow with HITL, something graph-based like LangGraph or a Temporal workflow. For a coding-style agent that needs file and shell tools, the Claude Agent SDK. For OpenAI-centric triage with handoffs, the Agents SDK. But I'd prototype the loop by hand first to understand the prompts, then adopt a framework for persistence and tracing, not for the loop itself."

15. Inter-agent protocols: MCP and A2A

MCP: Model Context Protocol

An open protocol, introduced by Anthropic in Nov 2024, for connecting AI applications to tools, resources and prompts exposed by servers. Anthropic 2024 Client–server over JSON-RPC (local stdio or streamable HTTP). Servers expose tools (callable functions), resources (readable data) and prompts (templates). The point is to solve the M×N integration problem once: any MCP client can use any MCP server. Spec versions are date-stamped, e.g. 2025-11-25. In Dec 2025 Anthropic donated MCP to the Linux Foundation's new Agentic AI Foundation (AAIF), co-founded with Block and OpenAI. MCP blog 2025

A2A: Agent2Agent Protocol

An open standard, created by Google (Apr 2025) and donated to the Linux Foundation, for agent-to-agent communication across frameworks and vendors. Agents interoperate without exposing internal memory, tools or logic. A2A docs Core concepts: an Agent Card (a JSON description of an agent's identity, skills, endpoint and auth, used for discovery), tasks with a lifecycle (submitted → working → input-required → completed/failed…), messages made of parts, and artifacts as outputs, with streaming and push notifications for long tasks. The A2A docs say explicitly that it complements MCP rather than replacing it, and that it is not a sub-agent or tool-call protocol.

MCPA2A
ConnectsAgent/app ↔ tools, data, promptsAgent ↔ agent (opaque peers)
Counterpart isA capability provider (stateless-ish function/resource)An autonomous agent with its own reasoning, possibly long-running
Unit of workTool call → resultTask with lifecycle, messages, artifacts
DiscoveryClient lists server's toolsAgent Card
Typical useGive your agent GitHub/Slack/DB accessYour procurement agent delegates to a supplier's quoting agent
Main risksTool poisoning / prompt injection via tool descriptions and outputs; over-broad credentialsTrust/auth across orgs; untrusted agent outputs; data leakage
May be out of date

Protocol governance and features moved quickly in 2025–2026: MCP spec revisions (auth, async tasks, elicitation), A2A versions, and foundation membership. Some 2026 reports say A2A has also moved under the AAIF umbrella; I couldn't confirm that from a primary source, so check a2a-protocol.org and the AAIF site for current status.

Interview angle

"MCP vs A2A?" One line: MCP is how an agent uses tools; A2A is how an agent talks to another agent it doesn't control. Then go into security. MCP tool descriptions and outputs are untrusted input (prompt injection, "tool poisoning"), so you need allowlisted servers, least-privilege credentials, and human confirmation for destructive tools. For A2A, cross-org auth and treating peer outputs as untrusted data.

Go deeper

16. Failure modes and mitigations

Why agents are hard, quantitatively: if each step succeeds independently with probability \(p\), an \(n\)-step task succeeds with probability \(p^n\). At \(p = 0.95\), 10 steps gives about 60% and 30 steps about 21%; at \(p = 0.99\), 30 steps gives about 74%. Real errors aren't independent and agents can recover, but the intuition holds: long horizons punish small per-step error rates, so verification and recovery matter more than any one step's accuracy.

$$ P(\text{success}) \approx p^{\,n} \qquad 0.95^{10}\approx 0.60,\; 0.95^{30}\approx 0.21,\; 0.99^{30}\approx 0.74 $$
Failure modeSymptomRoot causesMitigations
Infinite loops / repetitionSame tool call over and over; oscillating between two statesUnhelpful error messages; no progress signal; ambiguous done-criteriaMax steps; repeated-call detection; actionable tool errors; explicit finish tool; nudge injection
Tool-call errorsInvalid args, wrong tool, hallucinated tool namesPoor tool descriptions; overlapping tools; too many toolsSchema validation + error-as-observation; fewer, distinct tools; examples in descriptions; strict/structured tool modes; tool search for large catalogs
Compounding errorsEarly wrong assumption poisons everything downstreamNo verification between steps; \(p^n\)Verification gates; checkpoints and rollback; tests; critic with clean context
Goal driftAgent solves a different/easier problem; scope creepLong contexts dilute the original instruction; distracting observationsRe-state objective in a persistent todo/plan; periodic "am I on task?" checks; restate task in subagent specs
Context overflow / context rotForgetting earlier facts; degraded reasoning late in long runsHuge tool outputs; long histories; quality drops with lengthTruncate/paginate outputs; compaction; external notes; subagents for exploration (see Memory)
Premature completion"Done!" when it isn't; tests edited to passModel's done-judgment ≠ realityEnvironment-verified done (tests, checks); forbid editing tests; reviewer agent
Over-delegationSpawning 10 subagents for a simple question; telephone-game lossOrchestrator not taught to scale effortEffort-scaling rules in prompt; cap fan-out; single-agent fast path
Inter-agent misalignmentDuplicated work; incompatible outputs; agents ignoring each otherVague task specs; missing shared decisionsStructured specs with decisions_already_made; single writer; shared artifacts
Cost blowupsOne request costs 100× medianQuadratic history; loops; fan-out; no cachingBudgets per run/tenant; prompt caching; cheaper models for workers; alerting on cost p99
Prompt injection / unsafe actionsAgent follows instructions in a web page or tool output; exfiltrates dataUntrusted content in context + powerful toolsLeast privilege; sandbox; HITL on sensitive actions; separate "reader" from "actor"; egress controls
Non-reproducibilityCan't debug a failure that happened onceStochastic sampling; changing tools/dataFull traces (prompts, tool I/O, versions); record/replay; evals on traces
Interview angle

"How do you know your agent works?" Strong answer: (1) offline evals of end-to-end task success on a representative set, with both final-outcome and trajectory metrics (steps, tool errors, cost); (2) environment-based graders where possible (tests, DB state) and LLM judges with rubrics elsewhere; (3) tracing every run in production; (4) online metrics such as resolution rate, escalation rate, cost/latency percentiles; (5) regression evals on every prompt/model/tool change. Mention τ-bench-style evaluation of tool–agent–user interaction Yao+ 2024 and measuring consistency across repeated trials (pass^k), not just pass@1.

17. How to choose an architecture

Work through these questions in order. Each "no" pushes you toward a simpler design.

1. Can a single well-prompted call (with retrieval) do it? yes → single call 2. Are the steps known in advance? yes → prompt chain / pipeline (+ gates) 3. Do inputs fall into distinct categories needing different handling? yes → router → per-route chains 4. Are there independent subtasks known up front? yes → parallel sectioning (or voting for confidence) 5. Is there a clear quality criterion and does feedback help? yes → add evaluator–optimizer loop 6. Is the number/kind of steps unpredictable, needing tool feedback? yes → single agent (ReAct loop) with budgets 7. Is it long-horizon? yes → + persistent plan/todo, compaction, checkpoints 8. Is it breadth-first, read-heavy, larger than one context? yes → orchestrator + parallel read-only subagents 9. Does a specialist need to own the conversation for a branch? yes → handoffs 10. Must it survive crashes/human waits of hours–days? yes → durable execution (Temporal-style)
ScenarioRecommended architectureWhyWatch out for
Extract fields from invoicesSingle call + schema validation (+ retry)Fixed task; no tools neededDon't build an agent
Customer support (high volume)Router → per-intent chains; narrow tool agent for investigations; handoffs to specialists; HITL on refundsPredictable majority, long tail needs toolsPolicy compliance; cost per ticket
Blog post generationPipeline: outline → draft → evaluator–optimizer for style/factsKnown stages; clear criteriaEvaluator echo chamber
Deep research reportOrchestrator + parallel search subagents + citation verifierBreadth-first, read-heavy, exceeds one contextToken cost (≈15× chat); source quality
Bug fix in a repoSingle-threaded coding agent; read-only explorer subagents; reviewer with clean context; tests as verifierWrites must be coherent; reads parallelizeTest tampering; sandbox security
Large codebase migration (1000 files)Plan → per-file/per-module workers in isolated worktrees (mechanical, independent changes) → tests per unit → single integratorTruly independent writes when changes are mechanicalShared interfaces: decide them up front and pass them in specs
Voice receptionistStreaming single agent (or handoffs), few fast tools, filler speechLatency budget ≲ 1 sSupervisor hops; slow tools
Multi-day back-office process with approvalsDurable workflow (Temporal) with LLM steps as activities; HITL signals; sagasCrash-safety, long waits, side effectsNon-idempotent tools; versioning in-flight runs
Cross-company procurementYour agent + A2A to external agents; MCP for internal toolsOpaque peers need a protocolTrust, auth, untrusted outputs
Intuition

The default ladder is: single call → workflow → single agent → single agent + subagents for reading → (rarely) true multi-agent writing. Climb a rung only when an eval shows the lower rung failing on real inputs, and record the cost multiplier you accepted.

Go deeper

Interview question bank

1. What's the difference between a workflow and an agent, and how do you decide which to build?

A workflow orchestrates LLM calls and tools through predefined code paths. An agent lets the LLM decide its own sequence of tool calls and when to stop. The decision comes down to whether you can enumerate the steps for your real input distribution. If you can, a workflow is cheaper, faster, more predictable and easier to test, and patterns like chaining, routing, parallelization and evaluator–optimizer cover a lot of ground. Choose an agent when the number and kind of steps depend on what the agent discovers (debugging, research, investigations), when the environment gives feedback the model can use, and when the task's value justifies higher variance and cost. In practice most production systems are hybrids: a workflow skeleton with agentic nodes for the open-ended parts. I'd decide with evals: build the simplest version, measure its failures, and add autonomy only where it fixes them.

2. Walk me through the anatomy of an agent loop. Where do production concerns live?

The harness builds context (system prompt, tool schemas, history), calls the model, and gets back either a final answer or tool calls. It validates and permission-checks each call, executes it (in parallel if independent) with timeouts, converts results or errors into observations (truncated), appends them, checkpoints state, and loops. Production concerns live in the harness, not the model: budgets (steps, tokens, dollars, wall-clock), loop detection, policy/HITL gates for risky tools, retries for transient failures, error-as-observation for semantic failures, context compaction, checkpointing for resumability, and tracing of every prompt and tool I/O. The stop condition should be explicit, either a finish tool or no tool calls, and ideally backed by environment verification rather than the model's own claim.

3. Compare ReAct, plan-and-execute, ReWOO and LLMCompiler.

They differ in when reasoning happens relative to observations. ReAct interleaves think → act → observe every step: very adaptive, but one LLM call per action and quadratic context growth. Plan-and-execute writes a plan first, executes steps (often with a cheaper executor), and replans on failure, which saves strong-model calls and gives a reviewable plan but risks stale plans. ReWOO plans everything up front with placeholders for tool outputs and never re-reads observations during planning; it reported about 5× token efficiency on HotpotQA but can't adapt to surprises. LLMCompiler plans a DAG of tool calls and executes independent ones in parallel, reporting up to 3.7× lower latency and 6.7× lower cost than ReAct, with replanning support. In practice I'd use a hybrid: coarse plan, ReAct execution, parallel tool calls for independent steps, replan on surprise.

4. Back-of-envelope: an agent averages 25 steps, a 6k-token base prompt and adds 1.5k tokens per step. Roughly how many input tokens per task, and how would you cut it?

Total input ≈ \(nP + \Delta n(n-1)/2 = 25 \times 6\text{k} + 1.5\text{k} \times 300 = 150\text{k} + 450\text{k} \approx 600\text{k}\) tokens per task. At 10k tasks/day that's about 6B input tokens/day, so it matters. Levers: (1) prompt caching on the stable prefix and growing history, which typically discounts cached reads heavily; (2) shrink Δ by truncating, paginating or summarizing tool outputs, which usually dominate; (3) compaction, i.e. summarize older turns once past a threshold; (4) fewer steps through better tools (one tool that does the common 3-step sequence), parallel tool calls, or CodeAct-style multi-op steps; (5) offload exploration to subagents so bulky observations never enter the main thread; (6) route easy tasks to a cheaper model or a workflow.

5. Explain the Cognition vs Anthropic debate on multi-agent systems. Who's right?

Cognition (June 2025) argued against multi-agent systems: subagents don't share the parent's full context, and "actions carry implicit decisions", so parallel workers make conflicting assumptions (their Flappy Bird example). They recommended a single-threaded agent with history compression. Anthropic (a day later) showed an orchestrator with parallel subagents beating a single agent by 90.2% on an internal research eval, explained mostly by spending more tokens across parallel context windows, at about 15× chat token cost. Both are right for different task structures. Research is breadth-first and read-only, so it parallelizes. Coding produces one coherent artifact, so parallel writers conflict. The synthesis, which Cognition's April 2026 follow-up also lands on: keep writes single-threaded and use extra agents for intelligence (exploration, review, a "smart friend" model) rather than parallel actions. Controlled studies (e.g. Kim et al. 2025) also find that multi-agent helps on decomposable tasks and hurts on sequential ones.

6. Design a deep-research agent. What are the components and key decisions?

Orchestrator–worker. A lead agent clarifies scope, writes a research plan to memory (so it survives context limits), and decides effort based on query complexity: one agent for a factoid, several parallel subagents for a broad question. Each subagent gets a structured spec (objective, scope, sources to prefer or avoid, output schema, tool budget) and runs its own search/fetch/read loop with parallel tool calls, returning condensed findings with source URLs and noted gaps. Large content goes to an artifact store, passed by reference. The lead evaluates coverage, launches follow-up rounds if needed, synthesizes, and a citation pass verifies each claim against fetched text. Cross-cutting: per-run token budget split across children, dedup of sources, rate-limit semaphores, tracing. Evaluation uses LLM-judge rubrics (factuality, citation accuracy, completeness, source quality, efficiency) plus human review. Expect high token cost, so it's for high-value queries.

7. Supervisor vs handoffs: when would you use each?

Both route work among specialists, but they differ in who owns the conversation. With a supervisor (or agents-as-tools), a central agent calls specialists and gets control back each time, so it can synthesize across them and enforce policy in one place. The costs are an extra LLM hop per turn and a supervisor bottleneck. With handoffs, the active agent transfers control and the specialist talks to the user directly. That's lower latency and gives narrow prompts, but there's no global view, ping-pong is a risk, and it's weak at multi-domain requests that need synthesis. I'd use handoffs for conversational triage (support, voice) where a branch should be fully owned by a specialist, and a supervisor when the final answer combines multiple specialists' outputs or strict central policy is needed. In OpenAI's Agents SDK, handoffs appear to the model as transfer_to_<agent> tools, and history passed to the receiver can be trimmed with an input filter.

8. What do you put in a subagent's task specification, and why?

A subagent knows nothing beyond its prompt, so the spec has to carry the context that would otherwise be lost. It should include: a precise objective and why it matters (lets the subagent make sensible judgment calls); scope (include/exclude, sources); what's already known (to avoid duplicated work); decisions already made (interfaces, style, conventions), which directly addresses conflicting implicit decisions; constraints (tool-call/token budget, deadline, permissions); the output format (a schema, ideally); a done-criterion; and instructions to report uncertainty and gaps. Anthropic reported that vague one-line delegations caused subagents to duplicate work or miss coverage. Ask for condensed outputs or artifact references, not transcripts, to keep the parent's context clean.

9. If each step of an agent is 97% reliable, how reliable is a 20-step task? What does that imply for design?

Assuming independence, \(0.97^{20} \approx 0.54\), roughly a coin flip. At 40 steps it's about 0.30. Errors aren't truly independent and agents can recover, but the lesson holds: long horizons amplify small error rates. Design implications: shorten horizons (better tools that collapse multi-step sequences, workflows for known parts); add verification gates so errors are caught near where they happen (tests, schema checks, critics with clean context); make steps reversible or checkpointed so you can roll back; persist a plan to fight drift; and measure per-step error types in traces to target the dominant failure. Also evaluate consistency across repeated trials, not just single-run success.

10. How would you make a long-running agent survive crashes and multi-day human approvals?

Treat it as a durable workflow. In a Temporal-style design, the orchestration loop is deterministic workflow code, while every LLM call and tool execution is an activity whose result is recorded in event history. After a crash, a worker replays the history and reuses recorded results instead of re-calling the LLM (which would give different outputs). Human approvals arrive as signals, so the workflow can wait days without holding resources. Side-effecting tools take idempotency keys so retries are safe; transient errors get exponential-backoff retries, and semantic errors go back to the model. Multi-system side effects use sagas with registered compensations. Timeouts sit at every level, and workflow/prompt versions are pinned for in-flight runs. With LangGraph, persistent checkpointers plus interrupts give resumability, but you still have to make pre-interrupt side effects idempotent because the node re-runs on resume.

11. What's CodeAct and when would you prefer it to JSON tool calling?

CodeAct makes executable Python the agent's action space. Instead of one JSON tool call per turn, the model writes code that can call several tools, loop, branch, and keep intermediate values in variables, and it sees the printed output or traceback as the observation. The paper reported up to 20% higher success than JSON/text actions across 17 LLMs. Benefits: fewer turns (lower latency and context growth), natural composition, bulky intermediate data stays out of the context window, and it plays to models' coding strength. Downsides: you need a real sandbox (arbitrary code execution is a security boundary), it's harder to apply per-tool permissions and audit individual actions, and failures are messier. I'd prefer it for data-wrangling and multi-step tool composition in a sandbox, and keep typed JSON tools for sensitive, auditable actions like payments.

12. How do you prevent an agent from looping forever or blowing the budget?

Defense in depth in the harness: hard caps on steps, tokens/dollars and wall-clock time; detection of repeated identical tool calls or repeated errors, with a nudge injected or an abort; actionable tool error messages (a common cause of loops is an error the model can't act on); an explicit finish tool and clear done-criteria; observation truncation and compaction to limit quadratic growth; per-tenant and per-run budgets split across subagents; capped fan-out width; and alerting on cost and step-count percentiles. Offline, track the step-count distribution in evals so regressions show up before production. Prompt caching cuts the cost of the loops you do allow.

13. How would you design a coding agent? What matters most?

A single-threaded agent loop over a repo in a sandbox (container/microVM, no prod secrets, restricted egress). The tools are an agent–computer interface: search (grep/glob with concise output), a file viewer with windowing, an edit tool using exact search/replace (fails loudly on non-unique matches, so the agent must read before writing), a shell for tests/linters with output truncation, and git. Behaviors: reproduce the bug with a test first, make minimal edits, run tests, iterate on failures, present a diff. Planning lives in a todo list; long tasks keep progress files and commits. Subagents are used for read-only exploration and for a clean-context reviewer; writes stay in one thread. Guardrails: permission tiers (auto-run tests; approve network/installs/push), and checks against deleting or weakening tests. SWE-agent's main lesson is that ACI design moves results a lot, so invest in the tools.

14. What's the difference between MCP and A2A? Do you need both?

MCP standardizes how an AI application connects to tools, resources and prompts exposed by servers. It's agent-to-capability, JSON-RPC over stdio or HTTP, and solves the M×N integration problem. A2A standardizes communication between autonomous agents that are opaque to each other, possibly across vendors or organizations. It has Agent Cards for discovery, tasks with lifecycles, messages and artifacts, and streaming for long tasks. You need MCP whenever your agent uses external tools in a reusable way. You need A2A only when delegating to agents you don't control, such as a partner's agent. Within one codebase, subagents are usually just function calls or framework constructs, not A2A. Security differs too: MCP's main risk is untrusted tool descriptions/outputs (injection) and over-broad credentials; A2A's is cross-org trust and auth. MCP is now under the Linux Foundation's Agentic AI Foundation, and A2A was donated to the Linux Foundation by Google.

15. Your multi-agent system produces inconsistent outputs: subagents duplicate work and contradict each other. How do you debug and fix it?

First, trace it: look at the actual task specs the orchestrator sent and what each subagent returned. Most problems are specification failures at the boundary. Fixes: richer structured specs (objective, why, scope boundaries to prevent overlap, decisions already made, output schema); have the orchestrator partition the work explicitly (disjoint sources or questions) and track assignments in a ledger; route shared decisions (interfaces, style) through the orchestrator up front; switch from parallel writers to a single writer with parallel readers; add a synthesis/verification step that detects contradictions and resolves them with evidence; and if coupling is inherent, collapse back to a single agent with compaction. The MAST taxonomy is useful here because it separates system-design, inter-agent misalignment and verification failures.

16. When does planning actually help an agent, and how should the plan be represented?

Planning helps on long-horizon, decomposable tasks, on tasks where a human should approve the approach, and on tasks spanning multiple context windows. It adds overhead on short tasks and goes stale quickly in highly reactive environments. Representation should match the need: a structured todo list the agent re-reads and updates each turn (anti-drift, progress visibility), a DAG when you want parallel execution of independent steps, and plan/progress files plus git for multi-session work, since they survive context resets and are editable by humans. Key properties: each item has a verifiable done-criterion, explicit replan triggers (failure, contradiction, periodic review), and completion is verified by the environment rather than self-reported. Reasoning models plan implicitly, so explicit plans matter most for persistence and coordination.

17. How would you add human-in-the-loop to an agent without making it unusable?

Gate by risk, not by default. Classify tools by reversibility and blast radius: auto-allow reads and reversible actions, require approval for irreversible or external ones (payments, emails, deletions, deploys), and deny some outright. Implement approval as an interrupt: checkpoint state, show the reviewer the proposed action, arguments and the agent's rationale, then resume with approve/edit/reject; editing arguments is often more useful than a binary choice. Add policy thresholds (auto-approve refunds below a limit), timeouts with safe defaults, and batch approvals where possible. Make the gated actions idempotent so resumes can't double-execute. Track approval rates. If humans approve 99.9% without changes, the gate is either unnecessary or is being rubber-stamped, and both are problems.

18. Design a voice customer-service agent. How does the architecture differ from a text agent?

Latency dominates: responses need to start within roughly a second, so everything streams. Choose between a cascaded pipeline (VAD/endpointing → streaming STT → LLM → streaming TTS; modular and inspectable) and a speech-to-speech realtime model (lower latency, less control). Turn-taking needs good endpointing, barge-in handling (stop TTS, truncate the assistant message to what was spoken) and backchannels. Keep the agent graph shallow: a single agent or handoffs rather than a supervisor hop per turn. Use a few fast tools, filler speech or async execution for slow ones, and prefetch likely data (caller ID → account lookup). Prompts favor short, speakable sentences. Guardrails and HITL still apply to money actions, often via transfer to a human. Measure latency percentiles per stage, interruption rates and resolution rates.

19. Should you use a framework like LangGraph or roll your own agent loop?

The core loop is small, maybe a hundred lines on a provider SDK, and writing it yourself gives full visibility and lets you use new provider features (parallel tools, caching, thinking) immediately. What's expensive to build is everything around it: persistence/checkpointing, HITL interrupts, streaming, tracing, retries, durable execution. I'd pick a framework when control flow is graph-shaped and you need those features quickly (LangGraph for explicit state graphs, OpenAI Agents SDK for handoff-centric apps, Claude Agent SDK for file/shell-heavy agents). I'd roll my own for latency-critical or unusual control flow, or when a workflow engine like Temporal already provides durability. Either way, I'd require being able to see the exact prompt at every step, keep tools behind MCP or a clean interface so they're portable, and avoid deep framework lock-in in business logic.

20. What failure modes are specific to multi-agent systems compared with single agents?

Beyond the single-agent failures (loops, tool errors, drift), multi-agent systems add: context loss at boundaries (subagents missing information the parent had); conflicting implicit decisions between parallel writers; duplicated or uncovered work from vague task splits; over-delegation (spawning many agents for simple tasks); error amplification when there's no central verification; ping-pong between agents in handoff topologies; termination problems in peer-to-peer chats; supervisor bottlenecks and straggler latency; and cost multiplication (about 15× chat tokens in Anthropic's research system). Debugging is also harder because failures emerge from interactions. Mitigations are structured specs, single-writer designs, centralized verification, effort-scaling rules, capped fan-out, per-run budgets, and end-to-end tracing with agent/span IDs.

21. How would you evaluate an agent architecture change (e.g., moving from single agent to orchestrator + subagents)?

Build an eval set that reflects the real task distribution, including easy queries (where multi-agent may only add cost) and hard breadth-first ones. Measure end-to-end success with environment-based or rubric-based graders, plus trajectory metrics: steps, tool errors, tokens, dollars, latency p50/p95, and consistency across repeated runs (pass^k), because agents are stochastic. Run several trials per task to get confidence intervals. Compare against a strong single-agent baseline (good tools, compaction, same total budget), since much of multi-agent's gain can come from spending more tokens. Slice results by task type to find where it helps and hurts. Then do a staged rollout with online metrics and cost monitoring. If the gain only shows up on a slice, route only that slice to the expensive architecture.

22. What is a blackboard architecture and when is it a good fit for LLM agents?

Agents coordinate through a shared workspace instead of messaging each other. Each reads the current state, contributes when it can, and a controller (or the agents themselves) decides who acts next based on that state. In LLM systems the blackboard is a structured state object (LangGraph state with reducers is a version of this), a shared document, a database, or a repo. It fits incremental, opportunistic problem solving with many specialized contributors and loose coupling, and the shared state is fully inspectable. Downsides: concurrent writes need conflict handling (single writer per key, reducers, optimistic versioning), the board grows and needs curation, and agents can silently overwrite each other's implicit decisions. I'd use it with a typed schema, append-only or reducer-based updates, and a controller that serializes writes to contested fields.