Track B · AI product engineering

Working with AI Coding Tools

"How do you use AI coding tools?" is now a standard interview question. Interviewers aren't checking whether you've heard of Claude Code or Cursor. They want to know if you can direct a coding agent like a capable but fallible colleague: scope the work, give it context, make it verify its output, review the result, and keep it inside safe boundaries. This page covers the tool landscape, how the tools work, the configuration primitives that are now similar across vendors, the workflows practitioners use, team rollout and security, and what the productivity evidence really shows.

TL;DR: the 8–12 things to be able to say out loud

  • Three modes: autocomplete, chat, and agent (a loop that reads, edits, runs and verifies). Agents run in the terminal, the IDE, or cloud sandboxes that open PRs. The job has moved from pair programming toward delegation plus review.
  • Every coding agent is the same loop: gather context → act → verify → repeat, using read/search/edit/shell/web tools. The harness (tools, permissions, context management) matters as much as the model.
  • Context is the scarce resource. Quality drops as the window fills. Clear between tasks, compact deliberately, push research into subagents, keep always-loaded instructions short.
  • The extension stack has converged: memory file (CLAUDE.md / AGENTS.md / rules) → on-demand skills (SKILL.md, an open standard) → subagents → hooks → MCP → plugins. The ideas are the same everywhere; only the file names differ.
  • Memory files are advisory; hooks and permissions are enforced. Anything that must happen every time belongs in a hook.
  • Give the agent a check it can run (tests, typecheck, screenshot). Without one, you are the verification loop.
  • Explore → plan → implement → commit, using a spec or plan mode for multi-file work and TDD where possible. Skip the plan for one-line diffs.
  • Parallel agents run in worktrees or cloud sessions. Your review bandwidth becomes the bottleneck.
  • Failure modes to watch for: claiming it's done without evidence, tampering with tests, scope creep, hallucinated APIs and packages, giant diffs, silent assumptions.
  • Security: prompt injection, the "lethal trifecta", secrets, slopsquatting. Contain them with a sandbox and egress controls, least-privilege tokens, and human review.
  • Evidence is mixed: lab and field RCTs show roughly 21–56% gains, but METR 2025 measured a 19% slowdown for experts in mature repos, and its 2026 follow-up says selection effects make that design unreliable. DORA 2025 calls AI an amplifier.
  • Vibe coding vs agentic engineering: skipping code review is fine for throwaway prototypes. In production, you own every line.

1. The landscape (late 2026)

Autocomplete vs chat vs agent

Autocomplete ("tab")

A small, fast model predicts the next lines from nearby context. You review every suggestion as it appears. This was the original Copilot experience and is still the lowest-risk mode.

Chat

Ask a question, get an explanation or a snippet, paste it in. Context is whatever you attach. The model can't check its answer against your repo.

Agent (interactive)

The model loops with tools: searches, reads, edits, runs tests and the shell, reads the output and iterates. You supervise and course-correct. Examples: Claude Code, Codex CLI, Gemini CLI, Cursor's agent, Copilot agent mode.

Agent (background / cloud)

The same loop in a remote sandbox, started from an issue, chat or CLI. It returns a branch or PR, and you review the result rather than the process. Examples: Codex cloud, Copilot cloud agent, Claude Code in the cloud, Cursor Cloud Agents, Jules, Devin.

Intuition

Going from autocomplete to background agents trades review granularity for leverage. With autocomplete you review 3 lines at a time; with a background agent, a 400-line PR after the fact. The engineering work moves upstream into scoping, specs and verification harnesses, and downstream into fast, skeptical review. That's the shift from "pair programmer" to "delegate".

The main tools, by where they run

Terminal agents. Claude Code runs in the terminal and also in IDEs, a desktop app, the browser and Slack, with the same agent loop everywhere. Anthropic docs 2026 OpenAI Codex CLI is open source, with codex exec for non-interactive runs, MCP, skills and plugins, and codex cloud for remote tasks. OpenAI docs 2026 Gemini CLI is Apache-2.0, uses GEMINI.md for context, has a headless -p mode, MCP, and a GitHub Action. GitHub 2026 GitHub Copilot CLI has a -p mode and --allow-tool / --deny-tool flags. GitHub docs 2026 Aider is the veteran open-source terminal pair programmer. Aider 2026 You may also hear about Amp, OpenCode, Cursor CLI, Factory and goose.

IDE agents. Cursor is a VS Code-derived editor with Tab completion, an agent, rules, skills, hooks, subagents and MCP. Cursor docs 2026 Copilot in VS Code has agent mode, MCP, custom agents and checkpoints. VS Code docs 2026 Windsurf was acquired by Cognition in July 2025 and reportedly rebranded as Devin Desktop in June 2026. Wikipedia 2026 Cline is an open-source editor extension with Plan/Act modes, a CLI and MCP. Cline 2026 Roo Code, Kilo Code, JetBrains Junie and Kiro (spec-driven) are in the same category.

Cloud and background agents. Codex cloud runs a setup phase with network access, then an agent phase that is offline by default. OpenAI docs 2026 Copilot cloud agent (formerly "coding agent") picks up issue assignments, @copilot comments or chat prompts, works in an ephemeral GitHub Actions environment, and opens a draft PR. GitHub and Playwright MCP are enabled by default. GitHub docs 2026 Claude Code in the cloud runs on Anthropic-managed VMs. You start it from the browser, phone or claude --cloud, pull it back with --teleport, and it can auto-fix PRs. Anthropic docs 2026 The Claude Code GitHub Action responds to @claude mentions or runs prompts on GitHub events. Anthropic docs 2026 Cursor Cloud Agents run in VMs and can be launched from the editor, web, Slack, GitHub, Linear or the API. They return PRs with screenshots and logs. Cursor docs 2026 Jules clones your repo into a VM and proposes a plan first. Google 2026 Devin works in its own shell, IDE and browser. Cognition docs 2026

ToolInterfaceWhere code executesAutonomyExtensibility
Claude CodeTerminal; IDE, desktop, web, Slack, CILocal; cloud VMs; ActionsPermission modes (manual → auto); -p; cloudCLAUDE.md/AGENTS.md, skills, subagents, hooks, MCP, plugins, Agent SDK
CodexTerminal, IDE, ChatGPTLocal sandbox; cloudSandbox + approval policy; exec; cloudAGENTS.md, skills, plugins, MCP, hooks, TOML agents, SDK
Gemini CLITerminalLocal; ActionInteractive; -pGEMINI.md, skills, TOML commands, extensions, hooks, MCP
GitHub CopilotIDE, CLI, github.comLocal; Actions-based cloud agentAutocomplete → agent → cloud PRsInstructions files, AGENTS.md, prompt files, custom agents, skills, hooks, MCP
CursorIDE (+ CLI, web)Local; cloud VMsTab → agent → cloudRules, AGENTS.md, skills, subagents, hooks, MCP
Cline / AiderExtension, CLI / terminalLocalStep approval or autoRules/conventions, MCP (Cline), any model
Jules / DevinWeb, chat, CLI/APIVendor VMAsync plan → PRRepo instructions (AGENTS.md), integrations
May be out of date

Based on vendor docs as of October 2026. Product names and packaging change often (Copilot's "coding agent" became "cloud agent"; Windsurf became Devin Desktop). Categories (interface, execution location, autonomy, primitives) stay accurate longer than product facts.

Interview angle

"Which tools do you use?" really means "do you have a considered setup?" Name your primary tool and why ("terminal agent because it composes with shell, CI and worktrees"), mention a background agent for scoped tickets, and show you know the primitives carry over between tools. Don't come across as a fan of one vendor.

Go deeper

2. How coding agents work under the hood

Every tool above is an LLM in a tool-use loop plus a harness that supplies tools, manages context and enforces permissions. Anthropic describes three blended phases (gather context, take action, verify results) and calls the surrounding layer the "agentic harness". Anthropic docs 2026 For general agent theory, see B4 · Agent Architectures.

Your task + memory file, skills list, tool schemas 1. Gather context grep / glob / read files, git log, docs, MCP 2. Act edit files (search/replace), run shell commands 3. Verify tests, typecheck, lint, build, screenshot Observe tool output appended to context window Harness gates permissions sandbox (FS + net) hooks Done: diff + evidence stop when checks pass (or budget / human interrupt) loop: dozens of tool calls per task You can interrupt (Esc), redirect, or rewind at any point
The coding-agent loop. The model chooses each next action, the harness executes it under permission, sandbox and hook gates, and the result goes back into context. A runnable check (tests, build) is what lets the loop end on evidence rather than on "looks done".

Tools and edit formats

Claude Code groups its built-in tools as file operations, search (glob, regex), execution (shell, tests, git), web, and code intelligence (type errors and definitions via language-server plugins), plus orchestration tools for subagents and questions. Anthropic docs 2026 The shell is the key tool. It lets the agent use gh, test runners and any other CLI. Anthropic calls CLIs "the most context-efficient way to interact with external services". Anthropic docs 2026

How the model expresses an edit affects reliability and cost. Aider documents the options: Aider docs 2026

FormatProsCons
Whole fileSimple; nothing to matchExpensive; the model may silently drop code
Search/replace blocksCheap, precise, reviewableFails if the search text doesn't match exactly
Unified diffFamiliar; reduced "lazy" elisions for some modelsModels get hunks and line numbers wrong
Tool-call edit (path, old, new)Harness checks uniqueness and freshness, and snapshots for undoMany calls for sweeping changes

Modern harnesses mostly use the last option. The harness rejects an edit when old_string isn't unique or the file changed since the model last read it. Claude Code snapshots files before edits so you can rewind, but only for edits made through its file tools, not changes made via the shell. Anthropic docs 2026

Context gathering: agentic search vs embeddings

The trend is toward fast local search plus agentic exploration. Cursor's current docs describe a local Instant Grep index and say Cursor "does not store embeddings of your codebase for search", with an optional parallel Explore subagent. Cursor docs 2026 That's a change from its earlier embedding-based indexing.

Context windows and compaction

Anthropic's guide says most best practices follow from one constraint: the context window "fills up fast, and performance degrades as it fills". Anthropic docs 2026 Near the limit, Claude Code clears older tool outputs, then summarizes the conversation. Instructions given only in chat can be lost, while the root CLAUDE.md is re-read after compaction. Anthropic docs 2026 Anthropic docs 2026 In practice: filter big outputs (pytest -x -q, | tail -50) or let a subagent read them. Connected MCP tools used to cost context just by being present, and Claude Code now loads MCP tool definitions on demand by default ("tool search"). Anthropic docs 2026 A session full of failed attempts biases the model toward them, so restarting is often better than more corrections.

Permission modes and sandboxing

ToolMechanism (names as documented)
Claude CodeShift+Tab cycles modes: default (Manual), acceptEdits, plan, auto (a classifier blocks risky actions), plus dontAsk / bypassPermissions. Allow/deny rules in settings. An OS sandbox (/sandbox). Anthropic docs 2026
CodexSandbox read-only / workspace-write / danger-full-access, combined with an approval policy such as on-request or never. Network is off by default locally. OpenAI docs 2026
Copilot CLI--allow-tool, --deny-tool, --allow-all-tools, e.g. 'shell(rm)'. GitHub docs 2026
Cloud agentsA VM per task, egress allowlists, credentials held outside the sandbox behind a proxy. Anthropic docs 2026

Effective sandboxing needs both filesystem isolation (or the agent escapes) and network isolation (or it can exfiltrate keys). Claude Code uses Linux bubblewrap and macOS seatbelt, which also cover subprocesses. Anthropic reported 84% fewer permission prompts internally. Anthropic 2025

Common mistake

Running "YOLO mode" (bypassPermissions, danger-full-access, --allow-all-tools) on a laptop with production credentials in the environment. That's fine in a disposable, secret-free container with restricted egress. Anywhere else, you're one prompt injection away from an incident.

Model choice and reasoning effort

Harnesses let you pick the model and often a reasoning-effort level. Claude Code has /model, and Anthropic positions Sonnet for most coding and Opus for complex reasoning. Anthropic docs 2026 Codex custom agents can set model and model_reasoning_effort. OpenAI docs 2026 Rule of thumb: spend reasoning on planning and debugging, where one wrong turn wastes the session, and use cheaper models for mechanical fan-out and read-only search.

Interview angle

"Design a coding agent" (see B6): a minimal tool set; harness-validated exact-match edits; agentic search in a subagent; context budgeting and compaction; a verify step that runs tests; permission tiers plus an OS sandbox with egress control; checkpoints; and evals on real repo tasks to measure the harness separately from the model.

Go deeper

3. Configuration and extension primitives

Between 2025 and 2026 the major tools converged on the same extension points. Learn the concepts once and map the file names per tool. The layers run from "always in context, advisory" to "outside the model, enforced".

always loaded · advisory loaded on demand / outside the model Memory file CLAUDE.md · AGENTS.md · .cursor/rules · copilot-instructions.md · GEMINI.md: build cmds, conventions, gotchas Skills (SKILL.md) and slash commands only name + description loaded at start; body, scripts, references pulled in when relevant or invoked Subagents separate context window, own prompt / tools / model; return a summary to the parent Hooks and permission rules deterministic scripts at lifecycle events (PreToolUse, PostToolUse, Stop…): format, lint, block, gate MCP servers external tools and data: GitHub, Playwright browser, docs, DBs, issue trackers Plugins / marketplaces bundle skills + subagents + hooks + MCP config into one installable unit less context, more enforcement
The extension stack. Put a rule as low in the stack as it can go: stable facts in the memory file, occasional procedures in skills, noisy research in subagents, must-always rules in hooks or permissions, and external systems behind MCP or a CLI.

Project memory files

Markdown injected at session start, holding what you'd otherwise re-explain every time. Claude Code treats it as context, "not enforced configuration". Anthropic docs 2026

What makes a good one: test each line with "Would removing this cause Claude to make mistakes?" Include commands it can't guess, non-default style rules, test instructions, repo etiquette, architectural decisions and gotchas. Exclude what the code already shows, standard conventions, tutorials and file-by-file descriptions. Bloated files get ignored. Emphasize the one rule that keeps getting skipped, not every rule. Anthropic docs 2026

# AGENTS.md  (one file for every agent; CLAUDE.md just contains "@AGENTS.md")
## Commands
- Install: `pnpm install` (never npm/yarn)
- Typecheck: `pnpm tsc --noEmit`   Lint: `pnpm lint --fix`
- One test file: `pnpm vitest run path/to/file.test.ts` (prefer over full suite)
## Architecture (non-obvious bits)
- Handlers in `src/routes/*` call services in `src/services/*`; never import `src/db/*` directly.
- Money is integer cents (`type Cents = number`). Never floats.
## Conventions
- Throw `AppError(code, httpStatus)`; don't return `{ error }` objects.
- New endpoints need: zod schema + OpenAPI annotation + route test.
- Ask before adding any dependency.
## Gotchas
- Integration tests need `docker compose up -d db` (else ECONNREFUSED).
- `src/legacy/` is frozen. Don't refactor it.
## PRs
- Conventional commits; keep PRs under ~400 lines; description must say how you verified.
Common mistake

Treating the memory file as documentation (README, directory tree, API reference). It burns context every turn and buries the rules that matter. Link to docs instead, move procedures into skills and must-always rules into hooks, and add a line only when the agent repeats a mistake.

Slash commands and prompt files

These are named, reusable prompts. In Claude Code, commands have been merged into skills: .claude/commands/deploy.md and .claude/skills/deploy/SKILL.md both create /deploy, with $ARGUMENTS substitution. Anthropic docs 2026 Gemini CLI uses TOML in .gemini/commands/ (subfolders become /git:commit). Gemini CLI docs 2026 VS Code uses .github/prompts/*.prompt.md and recommends migrating to skills. VS Code docs 2026

Skills (Agent Skills, SKILL.md)

A skill is a folder with a SKILL.md (frontmatter plus instructions) and optional scripts/, references/ and assets/. Anthropic created the format and released it as an open standard. Supporting clients include Claude Code, Codex, Gemini CLI, Cursor, GitHub Copilot, VS Code, OpenCode, Goose, Junie and Kiro. agentskills.io 2026 The core idea is progressive disclosure: (1) only name + description load at startup (about 100 tokens); (2) the body loads when a task matches (recommended under 5,000 tokens and 500 lines); (3) scripts are executed and references read only when needed. Required fields: name (≤64 chars, lowercase/digits/hyphens, matches the folder) and description (≤1,024 chars, saying what the skill does and when to use it). Agent Skills spec 2026 Tools add fields of their own. Claude Code, for example, has disable-model-invocation: true for manual-only workflows. Anthropic docs 2026 Locations: .claude/skills/, .gemini/skills/ or .agents/skills/. Gemini CLI docs 2026

.claude/skills/db-migration/
├── SKILL.md
├── references/conventions.md      # expand/backfill/contract patterns, locking rules
└── scripts/check_migration.py     # lints a migration; run it, don't read it

---
name: db-migration
description: Create and validate PostgreSQL migrations for apps/api. Use when asked to
  add/alter/drop tables, columns or indexes, or when a migration is mentioned.
---
1. Generate: `pnpm db:new <snake_case_name>` (never hand-name files).
2. Forward-only SQL. Large tables: follow references/conventions.md.
   Never add NOT NULL without a default in the same migration.
3. Validate: `python .claude/skills/db-migration/scripts/check_migration.py <file>`.
   Fix every ERROR; justify each WARNING in the PR.
4. `pnpm db:migrate && pnpm vitest run src/db`. Paste both outputs in your summary.
Intuition

Memory file = what a new teammate needs on day one. Skill = the runbook they pull out for a specific task. The description is the highest-leverage line in a skill, because it's all the model sees when deciding whether to load it. Write it like a retrieval key, using the words users actually type.

Subagents

A subagent is a separately prompted agent with its own context window, tools and optional model that returns a summary. In Claude Code they're markdown files in .claude/agents/, requiring only name and description. They start without your conversation history, and the built-ins include read-only Explore and Plan. Anthropic docs 2026 Cursor reads .cursor/agents/ and also .claude/agents/ and .codex/agents/. Cursor docs 2026 Codex uses TOML files in .codex/agents/. OpenAI docs 2026

# .claude/agents/test-reviewer.md
---
name: test-reviewer
description: Reviews a diff for test quality after a feature or fix, before commit.
  Flags weakened assertions, skipped tests and special-casing.
tools: Read, Grep, Glob, Bash
model: sonnet
---
You are a skeptical reviewer. You see only the diff and the task, not the reasoning.
Report with file:line: (1) assertions loosened or deleted? (2) tests skipped or snapshots
regenerated? (3) implementation special-cases test inputs? (4) do new tests fail if the
implementation is reverted? Run them to confirm. Correctness issues only, no style nits.

Use subagents for: file-heavy research, so the noise stays out of the main context; independent review in a fresh context, so the agent that did the work isn't grading it Anthropic docs 2026; and parallel fan-out, optionally in worktrees (isolation: worktree). Anthropic docs 2026 Avoid them for tightly coupled edits. Subagents don't share context, so parallel writers make inconsistent decisions. Parallelize reading, keep writing single-threaded (B4).

Hooks

Hooks are commands (or HTTP calls, MCP tools, prompts) that run at lifecycle events. Memory files are advisory, while "hooks are deterministic and guarantee the action happens". Anthropic docs 2026 Claude Code events include SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop, SubagentStop and PreCompact. Exit code 2 blocks where the event supports it: PreToolUse blocks the call, and Stop keeps Claude working. PostToolUse can only feed stderr back. Anthropic docs 2026 The equivalents elsewhere: Cursor .cursor/hooks.json Cursor docs 2026, Copilot .github/hooks/*.json GitHub docs 2026, Codex .codex/hooks.json OpenAI docs 2026, and Gemini CLI. Gemini CLI docs 2026

// .claude/settings.json  (committed: same guardrails for everyone)
{
  "permissions": {
    "allow": ["Bash(pnpm lint *)", "Bash(pnpm vitest *)", "Bash(git diff *)"],
    "deny":  ["Read(./.env)", "Read(./.env.*)", "Bash(git push *)"]
  },
  "hooks": {
    "PreToolUse":  [{ "matcher": "Bash",
      "hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/block-dangerous.sh" }] }],
    "PostToolUse": [{ "matcher": "Edit|Write",
      "hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/format-and-lint.sh" }] }],
    "Stop":        [{
      "hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/require-green-typecheck.sh" }] }]
  }
}

# .claude/hooks/block-dangerous.sh  (tool call JSON arrives on stdin)
#!/usr/bin/env bash
cmd=$(jq -r '.tool_input.command // ""')
if echo "$cmd" | grep -Eq 'rm -rf /|git push --force|DROP TABLE|curl .*\| *(ba)?sh'; then
  echo "Blocked by policy: $cmd" >&2; exit 2     # 2 = block; stderr goes to Claude
fi

The allow/deny syntax follows the documented examples ("Bash(npm run test *)", "Read(./.env)"). Anthropic docs 2026 The Stop hook is the enforced version of "don't say done until typecheck passes". There's a cap on consecutive blocks so it can't loop forever. Anthropic docs 2026

Common mistake

Treating a regex denylist hook as a security boundary. Shell is too expressive (eval, base64, a Python one-liner). Hooks are for hygiene and policy. Containment comes from the sandbox and from which credentials exist at all.

MCP servers

MCP is the standard way to plug in external tools and data (protocol details in B4). MCP 2026 In Claude Code, servers are added with claude mcp add at local, project (.mcp.json, committed) or user scope. Anthropic docs 2026 Common choices: GitHub GitHub 2026 (though gh is cheaper on context); Playwright browser automation, so the agent can screenshot its own UI Microsoft 2026; up-to-date library docs (fewer hallucinated APIs); read-only databases; issue trackers; observability. Context cost: each tool's schema used to sit in every prompt, and many similar tools also hurt tool selection. Connect only what the project needs. Claude Code defers tool definitions via tool search by default and caps MCP output at 25,000 tokens unless configured. Anthropic docs 2026 A server that fetches external content is also an injection channel, so trust it as you would a dependency.

Plugins, headless mode and SDKs

Plugins bundle skills, hooks, subagents and MCP config into one installable unit (/plugin in Claude Code; skills become /plugin:skill). Anthropic docs 2026 Codex has plugins and Gemini CLI has extensions. GitHub 2026 A private marketplace is a natural way to share team defaults. Treat third-party plugins as code you run. Headless: claude -p with --output-format json and --allowedTools Anthropic docs 2026, codex exec OpenAI docs 2026, gemini -p Gemini CLI docs 2026, and copilot -p. SDKs: the Claude Agent SDK gives "the same tools, agent loop, and context management that power Claude Code" Anthropic docs 2026, and there's also the Codex SDK. OpenAI docs 2026 CI integrations are built on these.

Cross-tool mapping

PrimitiveClaude CodeCodexCursorGitHub CopilotGemini CLI
MemoryCLAUDE.md, .claude/rules/; reads AGENTS.mdAGENTS.md (+ override).cursor/rules/*.mdc; AGENTS.mdcopilot-instructions.md; *.instructions.md; AGENTS.mdGEMINI.md
Promptscommands → skillsskillscommands → skills*.prompt.md (VS Code).gemini/commands/*.toml
Skills.claude/skills/YesYesYes.gemini/ or .agents/skills/
Subagents.claude/agents/*.md.codex/agents/*.toml.cursor/agents/ (+ .claude/, .codex/)Custom agents (VS Code)check docs
Hookssettings.json.codex/hooks.json.cursor/hooks.json.github/hooks/*.jsonYes
MCP.mcp.jsoncodex mcpYesYesYes
Headlessclaude -p, Agent SDKcodex exec, SDKCursor CLIcopilot -pgemini -p

"Yes" means the docs confirm the feature exists, without our verifying the exact file path.

May be out of date

File names, event names and frontmatter keys change most often. Many tools added hooks, subagents and skills only in 2025–2026. Describe the primitive confidently and the file name loosely.

Interview angle

"CLAUDE.md, skill, hook or MCP: how do you decide?" Memory file for short, always-relevant facts. Skill for procedures needed sometimes. Subagent for context-heavy or independent work. Hook for anything that must always happen. MCP or a CLI for reaching external systems. Then add: hooks and permissions are guardrails, the sandbox is the boundary.

Go deeper

4. Daily workflows that work

These come mostly from vendor best-practice guides (Anthropic's is the most detailed) and from practitioners such as Simon Willison. The common thread: separate thinking from typing, and make the agent prove its work.

Explore → plan → implement → commit

Anthropic's recommended loop: (1) explore in plan mode (Shift+Tab or --permission-mode plan), where Claude reads and answers but doesn't edit; (2) plan, and edit the plan directly (Ctrl+G); (3) implement against the plan, writing and running tests; (4) commit and open a PR. Skip planning when "you could describe the diff in one sentence". Anthropic docs 2026 Cline's Plan/Act modes follow the same idea. Cline 2026

Spec-first development

For larger features, ask the agent to interview you about implementation, UX, edge cases and trade-offs, have it write SPEC.md, then start a fresh session to implement. Good specs are self-contained, name files and interfaces, state what's out of scope, and end with an end-to-end verification step. Anthropic docs 2026 GitHub's Spec Kit packages a constitution → specify → plan → tasks → implement workflow as agent skills. GitHub 2026 Kiro is built around specs. Kiro 2026

# SPEC.md: Rate limiting for /v1/* (skeleton)
Goal: per-API-key token bucket, 100 req/min, burst 20; 429 + Retry-After.
Non-goals: per-endpoint limits, UI, auth middleware changes.
Design: Redis bucket via Lua (src/infra/redis.ts); middleware src/middleware/rateLimit.ts after auth.
Edge cases: Redis down → fail open + warn + metric. Missing key → auth returns 401 first.
Acceptance (one test each): 1) burst of 120 → 20 succeed  2) keys isolated
  3) Retry-After integer ≥ 1  4) Redis down → requests succeed, warn logged.
Verify: `pnpm vitest run src/middleware` and `pnpm test:e2e -g rate` green; paste output.

TDD with agents

Tests are the ideal agent spec. Have the agent write tests from the acceptance criteria (no implementation, no mocking it), confirm they fail for the right reason, commit them, then implement until green without touching the tests. Anthropic also suggests one session writing tests and another writing the code. Anthropic docs 2026 Committing first matters because the most common agent cheat is editing the test, and a committed test makes that show up in the diff.

The verification loop

The habit with the most impact. Give the agent "a check it can run: tests, a build, a screenshot to compare". Without one, "you become the verification loop". Escalation: ask for the check in the prompt, then make it a session goal, then a Stop hook, then a fresh-context reviewer that tries to refute the result. Ask for evidence (command plus output, or a screenshot), not claims. Anthropic docs 2026

Change typeCheapest reliable check
Logic / libraryUnit tests written first; typechecker
API endpointRoute test against a local DB container; a curl script with expected output
UIBrowser-tool screenshot compared with the mock; visual-regression snapshot
Refactor / migrationSuite green before and after; typecheck; codemod dry-run count
Infra / configplan / --dry-run output reviewed by a human; never apply from the agent

UI iteration from mocks

Paste the mock, implement, screenshot the result, list the differences, fix, repeat. This is Anthropic's example prompt. Anthropic docs 2026 The agent needs a browser tool (for example Playwright MCP) to see its own output. Layout usually converges in a few rounds, but polish still needs a human eye.

Parallel agents with git worktrees

A worktree is a separate working directory on its own branch that shares the repo's history. One agent per worktree means edits never collide. Claude Code has this built in: claude --worktree feature-auth creates .claude/worktrees/feature-auth/ on branch worktree-feature-auth, .worktreeinclude copies gitignored files like .env, and you're prompted to clean up on exit. Plain git works with any tool. Anthropic docs 2026

main checkout shared .git you review here worktree A · feat/rate-limit agent 1: implement from SPEC.md worktree B · fix/flaky-auth-test agent 2: repro → fix worktree C · chore/upgrade-zod agent 3: migrate + typecheck tests green + evidence tests green + evidence stuck / asks a question You (human) review diffs answer Qs merge PRs = bottleneck git worktree add / claude --worktree · each agent has isolated files + branch · conflicts surface at merge, not mid-edit
Parallel agents in worktrees. Throughput is limited by how fast you can review and unblock, so pick tasks that are independent and come with their own checks.
# Plain git (any agent)
git worktree add ../app-ratelimit -b feat/rate-limit
git worktree add ../app-flaky     -b fix/flaky-auth-test
cp .env ../app-ratelimit/; cp .env ../app-flaky/          # gitignored files don't follow
(cd ../app-ratelimit && pnpm install && claude)            # terminal 1
(cd ../app-flaky     && pnpm install && codex)             # terminal 2
git worktree list;  git worktree remove ../app-flaky       # when merged

# Built in (Claude Code)
claude --worktree rate-limit      # .claude/worktrees/rate-limit, branch worktree-rate-limit
echo ".claude/worktrees/" >> .gitignore;  printf ".env\n" > .worktreeinclude

Parallel sessions also enable a writer/reviewer split. A fresh session isn't biased toward code it just wrote. Anthropic docs 2026 In practice, each worktree needs its own dependency install and its own ports, and most people's review quality drops beyond a handful of concurrent agents. That last point is folk wisdom, not a measured result.

Background agents for well-scoped tickets

Background agents are best for work that is well specified, independently verifiable and low-risk: a bug with a repro, a dependency bump, tests for an untested module, docs, a small feature behind a flag. One pattern: plan locally, commit the plan, then claude --cloud "Execute the migration plan in docs/migration-plan.md". Anthropic docs 2026 The cloud environment must be able to build and test the repo. Issue text written by anyone is an injection surface (§6).

Code review by and of agents

Agents make good first-pass reviewers when run in a fresh context with only the diff and the criteria. Because a reviewer told to find gaps "will usually report some, even when the work is sound", limit it to correctness and requirements. Anthropic docs 2026 When you review an agent's PR, read the tests first (meaningful? weakened?), then the core change, then scope and new dependencies. The output reads fluently whether it's right or not.

Debugging, onboarding, migrations, docs

Context hygiene and course-correction

A day in the life (illustrative)

08:45 Review 3 draft PRs from overnight background agents: merge 1, request changes 1, close 1 09:15 Plan mode on payments module; agent interviews me → SPEC.md; I edit non-goals; commit 09:45 Fresh session, worktree A: failing tests for criteria 1–4 → confirm they fail → commit 10:00 A: implement until green, tests untouched, typecheck + lint. Worktree B: flaky-test repro 10:40 Reviewer subagent vs SPEC.md flags missing Redis-down path → fixed. I read diff (tests first), PR 11:30 B: 3 hypotheses, experiment #2 confirms shared-fixture race → approve fix 14:30 Migration: agent writes codemod + list; dry-run 3 files; headless fan-out over 140; spot-check 10 16:45 Two repeated mistakes today → two new lines in AGENTS.md; queue 2 tickets for overnight
Interview angle

"Walk me through your day" is the most common question here. Give something like the timeline above: what you delegate and what you keep, how you spec, what verification you require, how you review, and one honest example of catching an agent mistake. Specifics persuade. "It 10x'd me" doesn't.

Go deeper

5. The skills that matter

Decomposition and scoping

Break work into pieces that are independently verifiable and small enough to review in one sitting. "Improve the dashboard" is a bad task. "Add server-side pagination to /orders with these 3 tests" is a good one.

Precise specs

Files, interfaces, constraints, non-goals, a verification command, an exemplar to follow. Every ambiguity becomes a silent assumption.

Context curation

A tight memory file, the right @-mentions, the real error, research pushed into subagents, clearing between tasks.

Delegate vs do

Delegate well-specified, verifiable, pattern-following work. Keep novel design, security-critical logic, subtle concurrency, and anything faster to write than to explain.

Fast, skeptical reading

Your throughput is your review speed. Read tests first, then the core change, then ripple effects. Look for what's missing.

Verification discipline

No "done" without evidence. Build the harness before delegating. If you can't verify it, don't ship it.

Security hygiene

Least-privilege credentials, sandboxes, no secrets in context, scrutiny of new dependencies, all external text treated as hostile.

Cost awareness

Know what drives cost (context size, model tier, effort, parallelism, retries) and use cheaper models for mechanical work.

Agent failure modes

Failure modeLooks likeCountermeasure
Overconfident completion"All tests pass ✅" after running a subset or noneRequire command + output; Stop hook; re-run important checks yourself
Test tampering / special-casingLoosened assertions, .skip, regenerated snapshots, if (input === testValue)Commit tests first; "don't modify tests"; review test diffs first; reviewer checklist
Scope creepUnrelated refactors and reformattingState non-goals; ask for a minimal diff; reject out-of-scope files
Hallucinated APIs / packagesPlausible methods, flags or packages that don't existTypecheck + tests; docs MCP; dependency allowlist and lockfile review
Giant diffs800 lines mixing feature, refactor, formattingDecompose; one concern per PR; split commits
Silent assumptionsPicks a timezone or retry policy without asking"List your assumptions"; questions before implementing; spec the edge cases
ThrashingRepeats the same failing fixInterrupt; after two failed corrections, clear and re-prompt
Common mistake

Reviewing agent code less carefully because it looks clean. LLM output lacks the surface signals that make reviewers slow down on weak human code. The errors are semantic: a wrong edge case, a wrong assumption, a test that asserts nothing.

Interview angle

"How do you know agent code is correct?" Answer in layers: before (spec, tests first), during (the agent runs checks, hooks enforce them), after (fresh-context reviewer, then human review starting with the tests), system (CI, branch protection, human approval). Close with: "I own it regardless of who typed it."

Go deeper

6. Team adoption and governance

Rolling coding agents out to a team is a change-management and security project, not a tooling purchase. DORA's 2025 report frames AI as an amplifier of existing strengths and dysfunctions, and finds that platform quality and clear AI policies determine whether individual gains show up at the organisation level (see §7). Google Cloud 2025

Shared configuration in the repo

Conventions for AI-authored PRs

Review standards

Volume is the risk: when PR counts double, rubber-stamping becomes the failure mode. Use an agent first pass for mechanical issues, a checklist for agent PRs (test changes, dependencies, scope, error paths), and CODEOWNERS-enforced scrutiny for auth, payments and infra.

Measuring impact (and the pitfalls)

MeasureWhyPitfall
Delivery metrics (lead time, deploy frequency, change failure rate, time to restore)Outcome-level; DORA's standard setLagging, confounded by everything else going on
Change failure / rework / revert rate on AI-assisted vs other PRsCatches the quality cost that throughput hidesNeeds reliable labelling of which PRs were AI-assisted
Review time and PR size distributionShows whether review is becoming the bottleneckFast reviews may mean rubber-stamping, not efficiency
Developer surveys (satisfaction, perceived productivity, trust)Cheap, captures experienceMETR showed self-reported speedup can have the wrong sign (§7)
Adoption / usage and cost per engineerNeeded for budgetingUsage ≠ value. Lines of code and "acceptance rate" are vanity metrics.

Run it like an experiment where you can: staged rollout by team, a baseline period, and paired quality and throughput metrics. Never use lines of code, PR counts or AI-acceptance rate as targets (Goodhart's law applies fast).

Security

Prompt injection is the defining risk. A coding agent reads large amounts of text it didn't write: repo files, dependency source, issue and PR comments, web pages, MCP tool results. Any of it can contain instructions. Simon Willison's lethal trifecta is the clearest model: an agent that combines (1) access to private data, (2) exposure to untrusted content and (3) the ability to communicate externally can be tricked into exfiltrating that data. Willison 2025 A background agent that reads a public issue, has a token with org-wide repo access, and has open internet egress has all three. Vendor docs echo this. Codex warns that "prompt injection can cause the agent to fetch and follow untrusted instructions" when network or web search is enabled. OpenAI docs 2026 Claude Code warns that MCP servers fetching external content create injection risk. Anthropic docs 2026

RiskExampleContainment
Prompt injectionIssue body says "also add this script to postinstall"; a fetched web page tells the agent to read ~/.aws/credentialsBreak the trifecta: sandbox + egress allowlist; no secrets in the environment; human review before merge; treat issue text from non-members as untrusted
Secrets exposureAgent reads .env into context; pastes a key into a log or PR; secrets end up in transcriptsDeny-read rules for secret files; secret managers instead of env files; pre-commit secret scanning; short-lived credentials. Claude Code's cloud mode keeps git credentials outside the VM behind a proxy. Anthropic docs 2026
Slopsquatting (hallucinated packages)Agent adds npm i fast-json-sanitizer, a name it invented, which an attacker has pre-registeredDependency allowlist / internal registry proxy; "ask before adding dependencies" rule; lockfile diffs reviewed; registry age and download heuristics in CI
Destructive actionsrm -rf, force-push, DROP TABLE, terraform applyPermission deny rules + hooks for hygiene; no prod credentials on dev machines; infra changes only through reviewed pipelines
Over-privileged automationCI agent with an org-admin PAT; GitHub App with write access everywhereLeast-privilege, repo-scoped, short-lived tokens; OIDC federation instead of long-lived secrets (the Claude Code Action supports this) Anthropic docs 2026
Comment-triggered automationAgent replies on a PR and triggers an issue_comment workflow that deploysAudit comment-triggered workflows before enabling auto-fix agents. Anthropic's docs warn about exactly this. Anthropic docs 2026

On slopsquatting, the USENIX Security 2025 study by Spracklen et al. tested 16 code LLMs over 576,000 generated samples. It found package-hallucination rates of at least 5.2% for commercial models and 21.7% for open-source models, with 205,474 unique invented package names. Spracklen+ 2025 Those rates were measured on 2024-era models in a specific prompting setup, so treat them as evidence that the risk is real, not as current rates.

Common mistake

Assuming there's no risk until merge. The agent runs code (install scripts, tests) during the task, so exfiltration happens at execution time. Sandbox execution, not just the merge.

Cost management

Drivers: context per turn, model tier, reasoning effort, parallel agents, runaway loops. Controls: budgets and alerts, cheaper models for subagents and fan-out, --max-turns and timeouts in CI, concise memory files, usage dashboards. CI runs consume both Actions minutes and tokens. Anthropic docs 2026 Anthropic docs 2026 Concrete per-seat prices change often, so look them up rather than quoting from memory.

Interview angle

"Roll out agents to a 50-person team?" Pilot with baseline metrics, then a paved road (shared config, sandbox defaults), a security review, PR conventions, enablement, paired throughput and quality measurement, and expansion team by team. Full model answer in the question bank. Flag that review capacity and CI speed become the new bottlenecks.

Go deeper

7. What the productivity evidence actually says

Expect to be asked "does AI actually make engineers faster?" The defensible answer is "it depends on the task, the developer, the codebase and the year, and the best evidence is mixed". Then show you know the studies.

StudyDesignHeadlineKey caveats
Peng et al. (GitHub/Microsoft), 2023 arXiv 2023Controlled experiment: implement an HTTP server in JavaScript, with vs without Copilot (autocomplete era)Treatment group finished ~55.8% fasterOne small greenfield task; a lab setting; authors affiliated with the vendor; pre-agent tooling
Cui, Demirer, Jaffe, Musolff, Peng, Salz (three field RCTs) MIT working paperRandomized Copilot access at Microsoft, Accenture and a Fortune 100 company; 4,867 developers~26% more completed tasks (SE ≈ 10 pts); larger gains and adoption among less-experienced developersOutcome proxies are PRs, commits and builds, not value or quality; autocomplete-era tool; imperfect compliance
Google enterprise RCT arXiv 202496 Google engineers, an enterprise-grade task, internal AI tooling, summer 2024~21% less time, with a wide confidence intervalA single lab task; authors caution the effect "will not necessarily apply more broadly"
METR, early-2025 METR 2025 arXiv 2025RCT: 16 experienced maintainers of large mature open-source repos, 246 real issues randomly assigned AI-allowed or not; mainly Cursor Pro with Claude 3.5/3.7 Sonnet19% slower with AI (CI roughly +2% to +39%). Developers predicted 24% faster and afterwards believed they'd been 20% fasterSmall n; experts in repos they know deeply (a setting where AI helps least); early-2025 tools; METR explicitly says it does not show AI fails to help most developers
METR follow-up (Feb 2026) METR 202657 developers (10 returning, 47 new), 143 repos, 800+ tasks, late-2025 toolsPoint estimates now lean toward speedup: about −18% time for returning devs (CI −38% to +9%) and −4% for new recruits (CI −15% to +9%)Both CIs include zero. Strong selection effects: developers declined to participate or withheld tasks they didn't want to do without AI, which biases the estimate. METR says the data "gives us an unreliable signal" and is changing its study design.
DORA 2025: State of AI-assisted Software Development Google Cloud 2025 DORA 2025Survey of nearly 5,000 professionals plus 100+ hours of qualitative data~90% use AI at work; 80%+ report higher productivity; ~30% have little or no trust in AI code. AI adoption now correlates positively with throughput but still negatively with delivery stability. AI acts as an "amplifier".Correlational, self-reported; describes associations, not causal effects; the capabilities model is DORA's framework
Intuition

Sort the studies by how much context the human already has. Greenfield tasks with less-experienced developers show big gains. Experts in codebases they know intimately, on issues full of tacit requirements, show little gain or a loss, because time saved typing goes into prompting, waiting and reviewing.

Common mistake

Citing vendor headline numbers ("55% faster!") as general truth, or citing METR's 19% as proof that "AI makes you slower". Both are over-generalizations of narrow studies. The METR perception gap is the most transferable finding: self-reported productivity is unreliable, so measure outcomes.

May be out of date

New studies appear every few months, and 2026 results with late-2025 agentic tools may already have superseded some of this table. Before an interview, search for the latest METR, DORA and peer-reviewed RCT results, and quote figures only with their caveats.

Interview angle

How to talk about it: "The controlled evidence ranges from big speedups on greenfield lab tasks to a measured slowdown for experts in mature repos, and METR's 2026 follow-up says selection effects now make that design unreliable. DORA's survey work suggests AI amplifies existing engineering practices: throughput up, stability not necessarily. So I treat it as a capability whose value depends on workflow. I measure outcomes rather than trusting how fast I feel, and I invest in what makes agents effective: tests, specs, fast CI and clear conventions." This shows you've read the research and won't overclaim.

Go deeper

8. Vibe coding vs agentic engineering

Vibe coding was coined by Andrej Karpathy in February 2025 for a style where you describe what you want, accept the AI's code without really reading it, and steer by running it and prompting again. Collins named it Word of the Year for 2025. Wikipedia 2026 Karpathy 2025 Simon Willison proposed vibe engineering (October 2025) for the opposite end: experienced engineers using agents to go faster while remaining accountable for the software, which leans heavily on tests, planning, documentation, version control, CI and code review. In a February 2026 update he noted that "agentic engineering" seemed to be becoming the more popular term. Willison 2025

Vibe codingAgentic engineering
Reads the code?Mostly no; judges by behaviourYes. Reviews every diff that ships
Verification"It seems to work when I click around"Tests, typecheck, CI, reviewer, evidence
SpecsConversational, emergentWritten specs, acceptance criteria, non-goals
Who's accountableNobody, reallyThe engineer, fully
Good forPrototypes, demos, personal tools, learning, exploring an idea, throwaway scriptsProduction systems, shared codebases, anything with users, data or money
Main riskSecurity holes, unmaintainable code, prototype becomes productionOver-process on trivial tasks; review fatigue

Most engineers do both, on purpose: vibe-code a prototype to test an idea, then rebuild it properly. The skill is noticing when a vibe-coded artifact is about to become production.

Interview angle

"What do you think about vibe coding?" Don't dismiss it and don't celebrate it. Say it's a legitimate tool for exploration and throwaway work and a liability for production, and that what matters is accountability: "if it ships under my name, I've read it, it's tested, and I can explain it." Then describe your own boundary, with an example of each.

Go deeper

Interview question bank

Walk me through how you use AI coding tools day to day.

Three modes: autocomplete while typing, an interactive terminal agent for feature work, debugging and codebase Q&A, and a background agent for well-scoped tickets I review asynchronously. For a feature, I explore in plan mode, have the agent interview me into a short spec with acceptance criteria and non-goals, then start a fresh session that writes failing tests first and implements until they pass, running typecheck and lint itself. A fresh-context reviewer checks the diff against the spec. Then I review it, tests first. Independent tasks run in parallel worktrees. Novel design and security-critical code I keep close. One real catch: an agent "fixed" a flaky test by adding a retry. Reading the test diff first exposed it, and we found the actual race.

How do you make sure agent-written code is correct?

Defense in depth. Before: a spec with acceptance criteria, and tests committed before implementation so any change to them is visible. During: the agent runs the real checks and pastes the evidence. For unattended runs, a Stop hook enforces it. After: an independent reviewer in a fresh context, then my review, starting with test changes (loosened assertions, skips, special-casing), then error paths and scope. System: CI, branch protection, required human approval. And I own it: if I can't explain a line, it doesn't ship.

How would you roll out coding agents to a 50-person engineering team?

Pilot first: 5–8 volunteers of mixed seniority on real work for 4–6 weeks, with a baseline captured before. Meanwhile build a paved road: a template AGENTS.md, shared skills and hooks, sandbox and permission defaults, approved MCP servers, all committed and reviewed like code. Run a security review (data retention, token scopes, egress, injection threat model, dependency policy). Set conventions: disclose AI-assisted PRs, include verification evidence, cap PR size, require human approval. Enable people with demos and a channel for sharing skills. Measure paired throughput and quality (lead time, change failure and revert rate, review time) plus cost, and distrust self-reports. Expand team by team, and expect review capacity and CI speed to become the bottlenecks.

When do you NOT use an agent?

When explaining the task takes longer than doing it. When I can't verify the result cheaply: subtle concurrency, numerics without reference outputs, security-critical auth. There I may use the agent to explore or draft, but I write and reason through the final code. When the work is mostly deciding rather than typing, though agents are useful sparring partners for that. When the task would touch production credentials and can't be sandboxed. And when I'm deliberately learning something, because writing it is how I learn.

How do you write a good CLAUDE.md / AGENTS.md?

Short, specific and checkable. Include what the agent can't infer: exact build, test and lint commands (including how to run one test), conventions that differ from defaults, architectural rules, environment gotchas, frozen areas, PR etiquette. Leave out what it can read from the code, generic advice, long docs (link them) and anything that changes often. The test for every line: would removing it cause mistakes? Keep it to a couple of hundred lines at most. Procedures go into skills, path-specific rules into scoped rule files, must-always rules into hooks. Use AGENTS.md as the cross-tool file and import it from CLAUDE.md. Review and prune it like code, and add a line when the agent repeats a mistake.

How do you manage context in a long session?

Performance degrades as the window fills, so: one task per session and clear between tasks. Push file-heavy research into subagents. Keep tool output small (filtered test runs, log tails). Keep durable state in files like PLAN.md so compaction or a fresh session loses nothing. Compact with a focus instruction, and put "what to preserve" in the memory file. After two failed corrections on the same issue, clear and re-prompt with what I learned, because failed attempts bias the model. When behaviour seems off, check what's loaded. Bloated memory files and many MCP tools eat budget silently.

What are the security risks of coding agents and how do you contain them?

The central one is prompt injection: the agent reads untrusted text (issues, dependency code, web pages, MCP results). Combined with private data and an outbound channel, that's the lethal trifecta. Others: secrets in context or logs, destructive commands, over-privileged CI tokens, slopsquatting, and comment-triggered workflows. Containment: a sandbox with filesystem and network isolation plus an egress allowlist; no long-lived secrets in the environment (proxy-brokered, short-lived, repo-scoped tokens, OIDC in CI); deny-read rules for secret files; a dependency allowlist and lockfile review; hooks for hygiene, while remembering they're not a boundary; and human review before merge. Code executes during the task, so isolate execution, not just merges.

Explain skills vs subagents vs hooks vs MCP. When would you use each?

Skill: a SKILL.md folder whose name and description load at startup, while the body and scripts load on demand. Use it for procedures needed sometimes. It's an open standard many tools read. Subagent: an agent with its own context, tools and model that returns a summary. Use it for context-heavy research, independent review or fan-out, not for coupled writes. Hook: a deterministic script at a lifecycle event. Use it for anything that must always happen (format, lint, block force-push, don't stop until typecheck passes). MCP: external tools and data. Use it when a CLI won't do, and watch context cost and injection risk. In short: knowledge in memory files and skills, enforcement in hooks and permissions, reach through MCP or a CLI.

How does a coding agent actually work under the hood?

An LLM in a tool-use loop. The harness sends the system prompt, memory files, skill descriptions, tool schemas and the conversation. The model emits tool calls (read, grep/glob, edit, shell, web). The harness checks permissions, runs them in a sandbox, appends results, and repeats until the model stops or a budget or interrupt ends it. Edits are usually exact-match replacements validated by the harness, with snapshots for undo. Context comes mostly from agentic search, often in a read-only subagent. When the window fills, old tool outputs are dropped and the conversation is summarized. Verification runs the project's own checks. The harness matters as much as the model.

An agent tells you "all tests pass." What do you do?

Ask for the command and its output, ideally already required by the prompt or a Stop hook. Then check the ways the claim can be true and still wrong. Did it run the whole relevant suite or one file? Did it modify, skip or loosen tests, or regenerate snapshots? Do the new tests fail if I revert the implementation? Is there special-casing of test inputs? For anything important, I re-run the check or rely on CI. A reviewer subagent with this checklist catches most of it cheaply.

How do you run multiple agents in parallel without chaos?

Isolation and independence. Each agent gets its own worktree, clone or cloud sandbox, with its own env files, dependencies and ports. I parallelize only tasks that are independent and come with their own checks. Coupled changes stay in one session, because parallel writers make inconsistent implicit decisions. Specs written up front mean agents rarely need me mid-task, and I batch reviews. The real limit is my review bandwidth: for me, 2–4 concurrent agents. A second session can also act as an unbiased reviewer of the first.

How do you use agents for a large refactor or migration?

If it's purely syntactic, use a deterministic codemod (ast-grep, jscodeshift, OpenRewrite), possibly written by the agent. If it needs judgment: have the agent produce a work list and a precise per-item prompt with a verification step. Run headless on 2–3 items, refine from the failures, then fan out with tightly scoped tool permissions, in parallel worktrees if supported. Each item reports OK or FAIL. Then run the full suite and typecheck, spot-check a random sample of diffs, handle failures by hand, and land the change in reviewable chunks.

What does the research say about AI coding productivity?

Mixed, and dependent on context. The 2023 Copilot lab study found about 56% faster completion of a small greenfield task. Field RCTs across three companies found about 26% more completed tasks, with bigger gains for juniors. Google's internal RCT found about 21% less time, with wide uncertainty. METR's 2025 RCT with expert open-source maintainers found a 19% slowdown, while participants believed they'd been 20% faster. Its 2026 follow-up leans toward speedup, but the CIs span zero and METR calls the signal unreliable because of selection effects. DORA 2025 sees AI as an amplifier: throughput up, stability not. So value depends on task, familiarity and workflow, and self-reports can't be trusted. Measure outcomes.

What's the difference between vibe coding and how you work?

Vibe coding (Karpathy's term) means accepting output without reading it and steering by behaviour. I use it deliberately for spikes, demos and throwaway tools. For production I do what Willison calls vibe engineering, now usually called agentic engineering: specs, tests first, the agent proving its work, and me reviewing every diff, because I'm accountable for what ships. The skill is choosing the mode on purpose and noticing when a prototype is about to become production.

How do you review a PR opened by a background agent?

With more suspicion than a colleague's PR, because it reads fluently whether or not it's right. Task and spec first, then test changes (do the new tests fail without the change, were old ones weakened?), then the core diff, then ripple effects. I check for scope creep, new dependencies (real? needed? reputable?), error handling and silent assumptions. I look at the attached evidence and re-run anything suspicious. If the approach is wrong, I close the PR and fix the spec rather than steer through comments. Usually the task wasn't specified well enough to delegate.

How do you keep costs under control?

The drivers are context size per turn, model tier and reasoning effort, parallel agents, and retries or loops. Keep memory files concise and MCP tools few, since they're paid for every turn. Use cheaper models for subagents, search and mechanical fan-out, and strong models for planning and hard debugging. Set turn limits and timeouts in CI, clear context instead of dragging long sessions along, and set budgets and dashboards per team. Track cost per merged, non-reverted change. A stronger model that gets it right first time is often cheaper overall.