Working with AI Coding Tools
"How do you use AI coding tools?" is now a standard interview question. Interviewers aren't checking whether you've heard of Claude Code or Cursor. They want to know if you can direct a coding agent like a capable but fallible colleague: scope the work, give it context, make it verify its output, review the result, and keep it inside safe boundaries. This page covers the tool landscape, how the tools work, the configuration primitives that are now similar across vendors, the workflows practitioners use, team rollout and security, and what the productivity evidence really shows.
TL;DR: the 8–12 things to be able to say out loud
- Three modes: autocomplete, chat, and agent (a loop that reads, edits, runs and verifies). Agents run in the terminal, the IDE, or cloud sandboxes that open PRs. The job has moved from pair programming toward delegation plus review.
- Every coding agent is the same loop: gather context → act → verify → repeat, using read/search/edit/shell/web tools. The harness (tools, permissions, context management) matters as much as the model.
- Context is the scarce resource. Quality drops as the window fills. Clear between tasks, compact deliberately, push research into subagents, keep always-loaded instructions short.
- The extension stack has converged: memory file (CLAUDE.md / AGENTS.md / rules) → on-demand skills (SKILL.md, an open standard) → subagents → hooks → MCP → plugins. The ideas are the same everywhere; only the file names differ.
- Memory files are advisory; hooks and permissions are enforced. Anything that must happen every time belongs in a hook.
- Give the agent a check it can run (tests, typecheck, screenshot). Without one, you are the verification loop.
- Explore → plan → implement → commit, using a spec or plan mode for multi-file work and TDD where possible. Skip the plan for one-line diffs.
- Parallel agents run in worktrees or cloud sessions. Your review bandwidth becomes the bottleneck.
- Failure modes to watch for: claiming it's done without evidence, tampering with tests, scope creep, hallucinated APIs and packages, giant diffs, silent assumptions.
- Security: prompt injection, the "lethal trifecta", secrets, slopsquatting. Contain them with a sandbox and egress controls, least-privilege tokens, and human review.
- Evidence is mixed: lab and field RCTs show roughly 21–56% gains, but METR 2025 measured a 19% slowdown for experts in mature repos, and its 2026 follow-up says selection effects make that design unreliable. DORA 2025 calls AI an amplifier.
- Vibe coding vs agentic engineering: skipping code review is fine for throwaway prototypes. In production, you own every line.
1. The landscape (late 2026)
Autocomplete vs chat vs agent
Autocomplete ("tab")
A small, fast model predicts the next lines from nearby context. You review every suggestion as it appears. This was the original Copilot experience and is still the lowest-risk mode.
Chat
Ask a question, get an explanation or a snippet, paste it in. Context is whatever you attach. The model can't check its answer against your repo.
Agent (interactive)
The model loops with tools: searches, reads, edits, runs tests and the shell, reads the output and iterates. You supervise and course-correct. Examples: Claude Code, Codex CLI, Gemini CLI, Cursor's agent, Copilot agent mode.
Agent (background / cloud)
The same loop in a remote sandbox, started from an issue, chat or CLI. It returns a branch or PR, and you review the result rather than the process. Examples: Codex cloud, Copilot cloud agent, Claude Code in the cloud, Cursor Cloud Agents, Jules, Devin.
Going from autocomplete to background agents trades review granularity for leverage. With autocomplete you review 3 lines at a time; with a background agent, a 400-line PR after the fact. The engineering work moves upstream into scoping, specs and verification harnesses, and downstream into fast, skeptical review. That's the shift from "pair programmer" to "delegate".
The main tools, by where they run
Terminal agents. Claude Code runs in the terminal and also in IDEs, a desktop app, the browser and Slack, with the same agent loop everywhere. Anthropic docs 2026 OpenAI Codex CLI is open source, with codex exec for non-interactive runs, MCP, skills and plugins, and codex cloud for remote tasks. OpenAI docs 2026 Gemini CLI is Apache-2.0, uses GEMINI.md for context, has a headless -p mode, MCP, and a GitHub Action. GitHub 2026 GitHub Copilot CLI has a -p mode and --allow-tool / --deny-tool flags. GitHub docs 2026 Aider is the veteran open-source terminal pair programmer. Aider 2026 You may also hear about Amp, OpenCode, Cursor CLI, Factory and goose.
IDE agents. Cursor is a VS Code-derived editor with Tab completion, an agent, rules, skills, hooks, subagents and MCP. Cursor docs 2026 Copilot in VS Code has agent mode, MCP, custom agents and checkpoints. VS Code docs 2026 Windsurf was acquired by Cognition in July 2025 and reportedly rebranded as Devin Desktop in June 2026. Wikipedia 2026 Cline is an open-source editor extension with Plan/Act modes, a CLI and MCP. Cline 2026 Roo Code, Kilo Code, JetBrains Junie and Kiro (spec-driven) are in the same category.
Cloud and background agents. Codex cloud runs a setup phase with network access, then an agent phase that is offline by default. OpenAI docs 2026 Copilot cloud agent (formerly "coding agent") picks up issue assignments, @copilot comments or chat prompts, works in an ephemeral GitHub Actions environment, and opens a draft PR. GitHub and Playwright MCP are enabled by default. GitHub docs 2026 Claude Code in the cloud runs on Anthropic-managed VMs. You start it from the browser, phone or claude --cloud, pull it back with --teleport, and it can auto-fix PRs. Anthropic docs 2026 The Claude Code GitHub Action responds to @claude mentions or runs prompts on GitHub events. Anthropic docs 2026 Cursor Cloud Agents run in VMs and can be launched from the editor, web, Slack, GitHub, Linear or the API. They return PRs with screenshots and logs. Cursor docs 2026 Jules clones your repo into a VM and proposes a plan first. Google 2026 Devin works in its own shell, IDE and browser. Cognition docs 2026
| Tool | Interface | Where code executes | Autonomy | Extensibility |
|---|---|---|---|---|
| Claude Code | Terminal; IDE, desktop, web, Slack, CI | Local; cloud VMs; Actions | Permission modes (manual → auto); -p; cloud | CLAUDE.md/AGENTS.md, skills, subagents, hooks, MCP, plugins, Agent SDK |
| Codex | Terminal, IDE, ChatGPT | Local sandbox; cloud | Sandbox + approval policy; exec; cloud | AGENTS.md, skills, plugins, MCP, hooks, TOML agents, SDK |
| Gemini CLI | Terminal | Local; Action | Interactive; -p | GEMINI.md, skills, TOML commands, extensions, hooks, MCP |
| GitHub Copilot | IDE, CLI, github.com | Local; Actions-based cloud agent | Autocomplete → agent → cloud PRs | Instructions files, AGENTS.md, prompt files, custom agents, skills, hooks, MCP |
| Cursor | IDE (+ CLI, web) | Local; cloud VMs | Tab → agent → cloud | Rules, AGENTS.md, skills, subagents, hooks, MCP |
| Cline / Aider | Extension, CLI / terminal | Local | Step approval or auto | Rules/conventions, MCP (Cline), any model |
| Jules / Devin | Web, chat, CLI/API | Vendor VM | Async plan → PR | Repo instructions (AGENTS.md), integrations |
Based on vendor docs as of October 2026. Product names and packaging change often (Copilot's "coding agent" became "cloud agent"; Windsurf became Devin Desktop). Categories (interface, execution location, autonomy, primitives) stay accurate longer than product facts.
"Which tools do you use?" really means "do you have a considered setup?" Name your primary tool and why ("terminal agent because it composes with shell, CI and worktrees"), mention a background agent for scoped tickets, and show you know the primitives carry over between tools. Don't come across as a fan of one vendor.
- How Claude Code works: loop, tools, environments.
- Codex CLI docs: exec mode and cloud handoff.
- About Copilot cloud agent: issue → draft PR, with guardrails.
2. How coding agents work under the hood
Every tool above is an LLM in a tool-use loop plus a harness that supplies tools, manages context and enforces permissions. Anthropic describes three blended phases (gather context, take action, verify results) and calls the surrounding layer the "agentic harness". Anthropic docs 2026 For general agent theory, see B4 · Agent Architectures.
Tools and edit formats
Claude Code groups its built-in tools as file operations, search (glob, regex), execution (shell, tests, git), web, and code intelligence (type errors and definitions via language-server plugins), plus orchestration tools for subagents and questions. Anthropic docs 2026 The shell is the key tool. It lets the agent use gh, test runners and any other CLI. Anthropic calls CLIs "the most context-efficient way to interact with external services". Anthropic docs 2026
How the model expresses an edit affects reliability and cost. Aider documents the options: Aider docs 2026
| Format | Pros | Cons |
|---|---|---|
| Whole file | Simple; nothing to match | Expensive; the model may silently drop code |
| Search/replace blocks | Cheap, precise, reviewable | Fails if the search text doesn't match exactly |
| Unified diff | Familiar; reduced "lazy" elisions for some models | Models get hunks and line numbers wrong |
Tool-call edit (path, old, new) | Harness checks uniqueness and freshness, and snapshots for undo | Many calls for sweeping changes |
Modern harnesses mostly use the last option. The harness rejects an edit when old_string isn't unique or the file changed since the model last read it. Claude Code snapshots files before edits so you can rewind, but only for edits made through its file tools, not changes made via the shell. Anthropic docs 2026
Context gathering: agentic search vs embeddings
- Agentic search: the model iteratively runs grep/glob/read and follows imports. It's always fresh, precise for identifiers, and needs no index. It costs more turns on vague questions. Claude Code works this way and uses a read-only
Exploresubagent to keep search noise out of the main context. Anthropic docs 2026 - Embeddings index: chunk, embed and retrieve by similarity. It helps when you don't know the identifier. The costs are staleness, chunking quality and privacy. This is RAG for code (see B2).
The trend is toward fast local search plus agentic exploration. Cursor's current docs describe a local Instant Grep index and say Cursor "does not store embeddings of your codebase for search", with an optional parallel Explore subagent. Cursor docs 2026 That's a change from its earlier embedding-based indexing.
Context windows and compaction
Anthropic's guide says most best practices follow from one constraint: the context window "fills up fast, and performance degrades as it fills". Anthropic docs 2026 Near the limit, Claude Code clears older tool outputs, then summarizes the conversation. Instructions given only in chat can be lost, while the root CLAUDE.md is re-read after compaction. Anthropic docs 2026 Anthropic docs 2026 In practice: filter big outputs (pytest -x -q, | tail -50) or let a subagent read them. Connected MCP tools used to cost context just by being present, and Claude Code now loads MCP tool definitions on demand by default ("tool search"). Anthropic docs 2026 A session full of failed attempts biases the model toward them, so restarting is often better than more corrections.
Permission modes and sandboxing
| Tool | Mechanism (names as documented) |
|---|---|
| Claude Code | Shift+Tab cycles modes: default (Manual), acceptEdits, plan, auto (a classifier blocks risky actions), plus dontAsk / bypassPermissions. Allow/deny rules in settings. An OS sandbox (/sandbox). Anthropic docs 2026 |
| Codex | Sandbox read-only / workspace-write / danger-full-access, combined with an approval policy such as on-request or never. Network is off by default locally. OpenAI docs 2026 |
| Copilot CLI | --allow-tool, --deny-tool, --allow-all-tools, e.g. 'shell(rm)'. GitHub docs 2026 |
| Cloud agents | A VM per task, egress allowlists, credentials held outside the sandbox behind a proxy. Anthropic docs 2026 |
Effective sandboxing needs both filesystem isolation (or the agent escapes) and network isolation (or it can exfiltrate keys). Claude Code uses Linux bubblewrap and macOS seatbelt, which also cover subprocesses. Anthropic reported 84% fewer permission prompts internally. Anthropic 2025
Running "YOLO mode" (bypassPermissions, danger-full-access, --allow-all-tools) on a laptop with production credentials in the environment. That's fine in a disposable, secret-free container with restricted egress. Anywhere else, you're one prompt injection away from an incident.
Model choice and reasoning effort
Harnesses let you pick the model and often a reasoning-effort level. Claude Code has /model, and Anthropic positions Sonnet for most coding and Opus for complex reasoning. Anthropic docs 2026 Codex custom agents can set model and model_reasoning_effort. OpenAI docs 2026 Rule of thumb: spend reasoning on planning and debugging, where one wrong turn wastes the session, and use cheaper models for mechanical fan-out and read-only search.
"Design a coding agent" (see B6): a minimal tool set; harness-validated exact-match edits; agentic search in a subagent; context budgeting and compaction; a verify step that runs tests; permission tiers plus an OS sandbox with egress control; checkpoints; and evals on real repo tasks to measure the harness separately from the model.
- Effective context engineering for AI agents: context as a finite budget.
- Claude Code sandboxing: why both FS and network isolation.
- Aider edit formats: edit-format trade-offs.
3. Configuration and extension primitives
Between 2025 and 2026 the major tools converged on the same extension points. Learn the concepts once and map the file names per tool. The layers run from "always in context, advisory" to "outside the model, enforced".
Project memory files
Markdown injected at session start, holding what you'd otherwise re-explain every time. Claude Code treats it as context, "not enforced configuration". Anthropic docs 2026
- CLAUDE.md: managed (org) → user
~/.claude/CLAUDE.md→ project./CLAUDE.md→ localCLAUDE.local.md(gitignored). Ancestor-directory files are concatenated at launch, and subdirectory files load when Claude reads files there.@pathimports pull in other files..claude/rules/*.mdwithpaths:globs load only for matching files. The docs suggest staying under about 200 lines. Anthropic docs 2026 - AGENTS.md: "a simple, open format for guiding coding agents", stewarded by the Agentic AI Foundation (Linux Foundation). Adopters include Codex, Jules, Cursor, Gemini CLI, Copilot's agent, Aider, Devin, Zed and VS Code. Rule: "the closest AGENTS.md to the edited file wins". agents.md 2026 Codex concatenates root-to-cwd files (with
AGENTS.override.mdtaking priority) up to 32 KiB by default (project_doc_max_bytes). OpenAI docs 2026 Claude Code reads AGENTS.md when no CLAUDE.md exists, or you can import it with@AGENTS.md. Anthropic docs 2026 - Cursor rules:
.cursor/rules/*.mdcwithalwaysApply/description/globs, giving Always, Intelligently, Specific Files or Manually. Cursor also has nested AGENTS.md, plus User and Team rules. Cursor docs 2026 - Copilot:
.github/copilot-instructions.md, path-specific.github/instructions/*.instructions.md, AGENTS.md anywhere, or a root CLAUDE.md or GEMINI.md. GitHub docs 2026 Gemini CLI: GEMINI.md. Gemini CLI docs 2026
What makes a good one: test each line with "Would removing this cause Claude to make mistakes?" Include commands it can't guess, non-default style rules, test instructions, repo etiquette, architectural decisions and gotchas. Exclude what the code already shows, standard conventions, tutorials and file-by-file descriptions. Bloated files get ignored. Emphasize the one rule that keeps getting skipped, not every rule. Anthropic docs 2026
# AGENTS.md (one file for every agent; CLAUDE.md just contains "@AGENTS.md")
## Commands
- Install: `pnpm install` (never npm/yarn)
- Typecheck: `pnpm tsc --noEmit` Lint: `pnpm lint --fix`
- One test file: `pnpm vitest run path/to/file.test.ts` (prefer over full suite)
## Architecture (non-obvious bits)
- Handlers in `src/routes/*` call services in `src/services/*`; never import `src/db/*` directly.
- Money is integer cents (`type Cents = number`). Never floats.
## Conventions
- Throw `AppError(code, httpStatus)`; don't return `{ error }` objects.
- New endpoints need: zod schema + OpenAPI annotation + route test.
- Ask before adding any dependency.
## Gotchas
- Integration tests need `docker compose up -d db` (else ECONNREFUSED).
- `src/legacy/` is frozen. Don't refactor it.
## PRs
- Conventional commits; keep PRs under ~400 lines; description must say how you verified.
Treating the memory file as documentation (README, directory tree, API reference). It burns context every turn and buries the rules that matter. Link to docs instead, move procedures into skills and must-always rules into hooks, and add a line only when the agent repeats a mistake.
Slash commands and prompt files
These are named, reusable prompts. In Claude Code, commands have been merged into skills: .claude/commands/deploy.md and .claude/skills/deploy/SKILL.md both create /deploy, with $ARGUMENTS substitution. Anthropic docs 2026 Gemini CLI uses TOML in .gemini/commands/ (subfolders become /git:commit). Gemini CLI docs 2026 VS Code uses .github/prompts/*.prompt.md and recommends migrating to skills. VS Code docs 2026
Skills (Agent Skills, SKILL.md)
A skill is a folder with a SKILL.md (frontmatter plus instructions) and optional scripts/, references/ and assets/. Anthropic created the format and released it as an open standard. Supporting clients include Claude Code, Codex, Gemini CLI, Cursor, GitHub Copilot, VS Code, OpenCode, Goose, Junie and Kiro. agentskills.io 2026 The core idea is progressive disclosure: (1) only name + description load at startup (about 100 tokens); (2) the body loads when a task matches (recommended under 5,000 tokens and 500 lines); (3) scripts are executed and references read only when needed. Required fields: name (≤64 chars, lowercase/digits/hyphens, matches the folder) and description (≤1,024 chars, saying what the skill does and when to use it). Agent Skills spec 2026 Tools add fields of their own. Claude Code, for example, has disable-model-invocation: true for manual-only workflows. Anthropic docs 2026 Locations: .claude/skills/, .gemini/skills/ or .agents/skills/. Gemini CLI docs 2026
.claude/skills/db-migration/
├── SKILL.md
├── references/conventions.md # expand/backfill/contract patterns, locking rules
└── scripts/check_migration.py # lints a migration; run it, don't read it
---
name: db-migration
description: Create and validate PostgreSQL migrations for apps/api. Use when asked to
add/alter/drop tables, columns or indexes, or when a migration is mentioned.
---
1. Generate: `pnpm db:new <snake_case_name>` (never hand-name files).
2. Forward-only SQL. Large tables: follow references/conventions.md.
Never add NOT NULL without a default in the same migration.
3. Validate: `python .claude/skills/db-migration/scripts/check_migration.py <file>`.
Fix every ERROR; justify each WARNING in the PR.
4. `pnpm db:migrate && pnpm vitest run src/db`. Paste both outputs in your summary.
Memory file = what a new teammate needs on day one. Skill = the runbook they pull out for a specific task. The description is the highest-leverage line in a skill, because it's all the model sees when deciding whether to load it. Write it like a retrieval key, using the words users actually type.
Subagents
A subagent is a separately prompted agent with its own context window, tools and optional model that returns a summary. In Claude Code they're markdown files in .claude/agents/, requiring only name and description. They start without your conversation history, and the built-ins include read-only Explore and Plan. Anthropic docs 2026 Cursor reads .cursor/agents/ and also .claude/agents/ and .codex/agents/. Cursor docs 2026 Codex uses TOML files in .codex/agents/. OpenAI docs 2026
# .claude/agents/test-reviewer.md
---
name: test-reviewer
description: Reviews a diff for test quality after a feature or fix, before commit.
Flags weakened assertions, skipped tests and special-casing.
tools: Read, Grep, Glob, Bash
model: sonnet
---
You are a skeptical reviewer. You see only the diff and the task, not the reasoning.
Report with file:line: (1) assertions loosened or deleted? (2) tests skipped or snapshots
regenerated? (3) implementation special-cases test inputs? (4) do new tests fail if the
implementation is reverted? Run them to confirm. Correctness issues only, no style nits.
Use subagents for: file-heavy research, so the noise stays out of the main context; independent review in a fresh context, so the agent that did the work isn't grading it Anthropic docs 2026; and parallel fan-out, optionally in worktrees (isolation: worktree). Anthropic docs 2026 Avoid them for tightly coupled edits. Subagents don't share context, so parallel writers make inconsistent decisions. Parallelize reading, keep writing single-threaded (B4).
Hooks
Hooks are commands (or HTTP calls, MCP tools, prompts) that run at lifecycle events. Memory files are advisory, while "hooks are deterministic and guarantee the action happens". Anthropic docs 2026 Claude Code events include SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop, SubagentStop and PreCompact. Exit code 2 blocks where the event supports it: PreToolUse blocks the call, and Stop keeps Claude working. PostToolUse can only feed stderr back. Anthropic docs 2026 The equivalents elsewhere: Cursor .cursor/hooks.json Cursor docs 2026, Copilot .github/hooks/*.json GitHub docs 2026, Codex .codex/hooks.json OpenAI docs 2026, and Gemini CLI. Gemini CLI docs 2026
// .claude/settings.json (committed: same guardrails for everyone)
{
"permissions": {
"allow": ["Bash(pnpm lint *)", "Bash(pnpm vitest *)", "Bash(git diff *)"],
"deny": ["Read(./.env)", "Read(./.env.*)", "Bash(git push *)"]
},
"hooks": {
"PreToolUse": [{ "matcher": "Bash",
"hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/block-dangerous.sh" }] }],
"PostToolUse": [{ "matcher": "Edit|Write",
"hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/format-and-lint.sh" }] }],
"Stop": [{
"hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/require-green-typecheck.sh" }] }]
}
}
# .claude/hooks/block-dangerous.sh (tool call JSON arrives on stdin)
#!/usr/bin/env bash
cmd=$(jq -r '.tool_input.command // ""')
if echo "$cmd" | grep -Eq 'rm -rf /|git push --force|DROP TABLE|curl .*\| *(ba)?sh'; then
echo "Blocked by policy: $cmd" >&2; exit 2 # 2 = block; stderr goes to Claude
fi
The allow/deny syntax follows the documented examples ("Bash(npm run test *)", "Read(./.env)"). Anthropic docs 2026 The Stop hook is the enforced version of "don't say done until typecheck passes". There's a cap on consecutive blocks so it can't loop forever. Anthropic docs 2026
Treating a regex denylist hook as a security boundary. Shell is too expressive (eval, base64, a Python one-liner). Hooks are for hygiene and policy. Containment comes from the sandbox and from which credentials exist at all.
MCP servers
MCP is the standard way to plug in external tools and data (protocol details in B4). MCP 2026 In Claude Code, servers are added with claude mcp add at local, project (.mcp.json, committed) or user scope. Anthropic docs 2026 Common choices: GitHub GitHub 2026 (though gh is cheaper on context); Playwright browser automation, so the agent can screenshot its own UI Microsoft 2026; up-to-date library docs (fewer hallucinated APIs); read-only databases; issue trackers; observability. Context cost: each tool's schema used to sit in every prompt, and many similar tools also hurt tool selection. Connect only what the project needs. Claude Code defers tool definitions via tool search by default and caps MCP output at 25,000 tokens unless configured. Anthropic docs 2026 A server that fetches external content is also an injection channel, so trust it as you would a dependency.
Plugins, headless mode and SDKs
Plugins bundle skills, hooks, subagents and MCP config into one installable unit (/plugin in Claude Code; skills become /plugin:skill). Anthropic docs 2026 Codex has plugins and Gemini CLI has extensions. GitHub 2026 A private marketplace is a natural way to share team defaults. Treat third-party plugins as code you run. Headless: claude -p with --output-format json and --allowedTools Anthropic docs 2026, codex exec OpenAI docs 2026, gemini -p Gemini CLI docs 2026, and copilot -p. SDKs: the Claude Agent SDK gives "the same tools, agent loop, and context management that power Claude Code" Anthropic docs 2026, and there's also the Codex SDK. OpenAI docs 2026 CI integrations are built on these.
Cross-tool mapping
| Primitive | Claude Code | Codex | Cursor | GitHub Copilot | Gemini CLI |
|---|---|---|---|---|---|
| Memory | CLAUDE.md, .claude/rules/; reads AGENTS.md | AGENTS.md (+ override) | .cursor/rules/*.mdc; AGENTS.md | copilot-instructions.md; *.instructions.md; AGENTS.md | GEMINI.md |
| Prompts | commands → skills | skills | commands → skills | *.prompt.md (VS Code) | .gemini/commands/*.toml |
| Skills | .claude/skills/ | Yes | Yes | Yes | .gemini/ or .agents/skills/ |
| Subagents | .claude/agents/*.md | .codex/agents/*.toml | .cursor/agents/ (+ .claude/, .codex/) | Custom agents (VS Code) | check docs |
| Hooks | settings.json | .codex/hooks.json | .cursor/hooks.json | .github/hooks/*.json | Yes |
| MCP | .mcp.json | codex mcp | Yes | Yes | Yes |
| Headless | claude -p, Agent SDK | codex exec, SDK | Cursor CLI | copilot -p | gemini -p |
"Yes" means the docs confirm the feature exists, without our verifying the exact file path.
File names, event names and frontmatter keys change most often. Many tools added hooks, subagents and skills only in 2025–2026. Describe the primitive confidently and the file name loosely.
"CLAUDE.md, skill, hook or MCP: how do you decide?" Memory file for short, always-relevant facts. Skill for procedures needed sometimes. Subagent for context-heavy or independent work. Hook for anything that must always happen. MCP or a CLI for reaching external systems. Then add: hooks and permissions are guardrails, the sandbox is the boundary.
- How Claude remembers your project: scopes, rules, AGENTS.md interop.
- Agent Skills specification: the open SKILL.md format.
- Claude Code hooks reference: events and exit codes.
- AGENTS.md: the cross-tool convention.
4. Daily workflows that work
These come mostly from vendor best-practice guides (Anthropic's is the most detailed) and from practitioners such as Simon Willison. The common thread: separate thinking from typing, and make the agent prove its work.
Explore → plan → implement → commit
Anthropic's recommended loop: (1) explore in plan mode (Shift+Tab or --permission-mode plan), where Claude reads and answers but doesn't edit; (2) plan, and edit the plan directly (Ctrl+G); (3) implement against the plan, writing and running tests; (4) commit and open a PR. Skip planning when "you could describe the diff in one sentence". Anthropic docs 2026 Cline's Plan/Act modes follow the same idea. Cline 2026
Spec-first development
For larger features, ask the agent to interview you about implementation, UX, edge cases and trade-offs, have it write SPEC.md, then start a fresh session to implement. Good specs are self-contained, name files and interfaces, state what's out of scope, and end with an end-to-end verification step. Anthropic docs 2026 GitHub's Spec Kit packages a constitution → specify → plan → tasks → implement workflow as agent skills. GitHub 2026 Kiro is built around specs. Kiro 2026
# SPEC.md: Rate limiting for /v1/* (skeleton)
Goal: per-API-key token bucket, 100 req/min, burst 20; 429 + Retry-After.
Non-goals: per-endpoint limits, UI, auth middleware changes.
Design: Redis bucket via Lua (src/infra/redis.ts); middleware src/middleware/rateLimit.ts after auth.
Edge cases: Redis down → fail open + warn + metric. Missing key → auth returns 401 first.
Acceptance (one test each): 1) burst of 120 → 20 succeed 2) keys isolated
3) Retry-After integer ≥ 1 4) Redis down → requests succeed, warn logged.
Verify: `pnpm vitest run src/middleware` and `pnpm test:e2e -g rate` green; paste output.
TDD with agents
Tests are the ideal agent spec. Have the agent write tests from the acceptance criteria (no implementation, no mocking it), confirm they fail for the right reason, commit them, then implement until green without touching the tests. Anthropic also suggests one session writing tests and another writing the code. Anthropic docs 2026 Committing first matters because the most common agent cheat is editing the test, and a committed test makes that show up in the diff.
The verification loop
The habit with the most impact. Give the agent "a check it can run: tests, a build, a screenshot to compare". Without one, "you become the verification loop". Escalation: ask for the check in the prompt, then make it a session goal, then a Stop hook, then a fresh-context reviewer that tries to refute the result. Ask for evidence (command plus output, or a screenshot), not claims. Anthropic docs 2026
| Change type | Cheapest reliable check |
|---|---|
| Logic / library | Unit tests written first; typechecker |
| API endpoint | Route test against a local DB container; a curl script with expected output |
| UI | Browser-tool screenshot compared with the mock; visual-regression snapshot |
| Refactor / migration | Suite green before and after; typecheck; codemod dry-run count |
| Infra / config | plan / --dry-run output reviewed by a human; never apply from the agent |
UI iteration from mocks
Paste the mock, implement, screenshot the result, list the differences, fix, repeat. This is Anthropic's example prompt. Anthropic docs 2026 The agent needs a browser tool (for example Playwright MCP) to see its own output. Layout usually converges in a few rounds, but polish still needs a human eye.
Parallel agents with git worktrees
A worktree is a separate working directory on its own branch that shares the repo's history. One agent per worktree means edits never collide. Claude Code has this built in: claude --worktree feature-auth creates .claude/worktrees/feature-auth/ on branch worktree-feature-auth, .worktreeinclude copies gitignored files like .env, and you're prompted to clean up on exit. Plain git works with any tool. Anthropic docs 2026
# Plain git (any agent)
git worktree add ../app-ratelimit -b feat/rate-limit
git worktree add ../app-flaky -b fix/flaky-auth-test
cp .env ../app-ratelimit/; cp .env ../app-flaky/ # gitignored files don't follow
(cd ../app-ratelimit && pnpm install && claude) # terminal 1
(cd ../app-flaky && pnpm install && codex) # terminal 2
git worktree list; git worktree remove ../app-flaky # when merged
# Built in (Claude Code)
claude --worktree rate-limit # .claude/worktrees/rate-limit, branch worktree-rate-limit
echo ".claude/worktrees/" >> .gitignore; printf ".env\n" > .worktreeinclude
Parallel sessions also enable a writer/reviewer split. A fresh session isn't biased toward code it just wrote. Anthropic docs 2026 In practice, each worktree needs its own dependency install and its own ports, and most people's review quality drops beyond a handful of concurrent agents. That last point is folk wisdom, not a measured result.
Background agents for well-scoped tickets
Background agents are best for work that is well specified, independently verifiable and low-risk: a bug with a repro, a dependency bump, tests for an untested module, docs, a small feature behind a flag. One pattern: plan locally, commit the plan, then claude --cloud "Execute the migration plan in docs/migration-plan.md". Anthropic docs 2026 The cloud environment must be able to build and test the repo. Issue text written by anyone is an injection surface (§6).
Code review by and of agents
Agents make good first-pass reviewers when run in a fresh context with only the diff and the criteria. Because a reviewer told to find gaps "will usually report some, even when the work is sound", limit it to correctness and requirements. Anthropic docs 2026 When you review an agent's PR, read the tests first (meaningful? weakened?), then the core change, then scope and new dependencies. The output reads fluently whether it's right or not.
Debugging, onboarding, migrations, docs
- Debugging: pipe in the real error (
cat error.log | claude), give the symptom, the likely location and what "fixed" means, ask for a failing reproduction test first, and say "address the root cause, don't suppress the error". Anthropic docs 2026 For murky bugs, ask for ranked hypotheses with a cheap experiment for each, because agents lock onto their first theory. - Onboarding Q&A: ask what you'd ask a senior engineer, for example "look through ExecutionFactory's git history and summarize how its API came to be". Anthropic docs 2026 Nothing is written, so the only risk is being misinformed. Ask for file:line citations.
- Migrations (fan-out): the agent produces a file list, a script loops
claude -p "Migrate $file ... Return OK or FAIL." --allowedTools "Edit,Bash(git commit *)", and you test on 2–3 files before running everything. Claude Code also has/batch, which splits a change across subagents in worktrees. Anthropic docs 2026 Use deterministic codemods for purely syntactic changes. - Docs, tests, scripts: good to delegate. Derive tests from the spec, not the implementation, or they'll lock in its bugs.
- Prototype vs production: decide which you're doing before starting. The failure is a prototype quietly becoming production (§8).
Context hygiene and course-correction
/clearbetween unrelated tasks to avoid "the kitchen sink session". Compact with a focus (/compact Focus on the API changes). Scope investigations or hand them to subagents. Anthropic docs 2026- Keep state in files (
PLAN.md, a checklist) so a fresh or compacted session can resume from disk. - Interrupt early (Esc). Rewind with Esc Esc or
/rewind. "If you've corrected Claude more than twice on the same issue", clear and re-prompt. Checkpoints don't cover shell side effects and don't replace git, so commit at known-good points. Anthropic docs 2026
A day in the life (illustrative)
"Walk me through your day" is the most common question here. Give something like the timeline above: what you delegate and what you keep, how you spec, what verification you require, how you review, and one honest example of catching an agent mistake. Specifics persuade. "It 10x'd me" doesn't.
- Best practices for Claude Code: the most detailed workflow guide; mostly tool-agnostic.
- Parallel sessions with worktrees: isolation and cleanup.
- GitHub Spec Kit: an opinionated spec-driven workflow.
5. The skills that matter
Decomposition and scoping
Break work into pieces that are independently verifiable and small enough to review in one sitting. "Improve the dashboard" is a bad task. "Add server-side pagination to /orders with these 3 tests" is a good one.
Precise specs
Files, interfaces, constraints, non-goals, a verification command, an exemplar to follow. Every ambiguity becomes a silent assumption.
Context curation
A tight memory file, the right @-mentions, the real error, research pushed into subagents, clearing between tasks.
Delegate vs do
Delegate well-specified, verifiable, pattern-following work. Keep novel design, security-critical logic, subtle concurrency, and anything faster to write than to explain.
Fast, skeptical reading
Your throughput is your review speed. Read tests first, then the core change, then ripple effects. Look for what's missing.
Verification discipline
No "done" without evidence. Build the harness before delegating. If you can't verify it, don't ship it.
Security hygiene
Least-privilege credentials, sandboxes, no secrets in context, scrutiny of new dependencies, all external text treated as hostile.
Cost awareness
Know what drives cost (context size, model tier, effort, parallelism, retries) and use cheaper models for mechanical work.
Agent failure modes
| Failure mode | Looks like | Countermeasure |
|---|---|---|
| Overconfident completion | "All tests pass ✅" after running a subset or none | Require command + output; Stop hook; re-run important checks yourself |
| Test tampering / special-casing | Loosened assertions, .skip, regenerated snapshots, if (input === testValue) | Commit tests first; "don't modify tests"; review test diffs first; reviewer checklist |
| Scope creep | Unrelated refactors and reformatting | State non-goals; ask for a minimal diff; reject out-of-scope files |
| Hallucinated APIs / packages | Plausible methods, flags or packages that don't exist | Typecheck + tests; docs MCP; dependency allowlist and lockfile review |
| Giant diffs | 800 lines mixing feature, refactor, formatting | Decompose; one concern per PR; split commits |
| Silent assumptions | Picks a timezone or retry policy without asking | "List your assumptions"; questions before implementing; spec the edge cases |
| Thrashing | Repeats the same failing fix | Interrupt; after two failed corrections, clear and re-prompt |
Reviewing agent code less carefully because it looks clean. LLM output lacks the surface signals that make reviewers slow down on weak human code. The errors are semantic: a wrong edge case, a wrong assumption, a test that asserts nothing.
"How do you know agent code is correct?" Answer in layers: before (spec, tests first), during (the agent runs checks, hooks enforce them), after (fresh-context reviewer, then human review starting with the tests), system (CI, branch protection, human approval). Close with: "I own it regardless of who typed it."
- Avoid common failure patterns: kitchen-sink sessions, over-correction.
- Vibe engineering: why senior skills make agents useful.
6. Team adoption and governance
Rolling coding agents out to a team is a change-management and security project, not a tooling purchase. DORA's 2025 report frames AI as an amplifier of existing strengths and dysfunctions, and finds that platform quality and clear AI policies determine whether individual gains show up at the organisation level (see §7). Google Cloud 2025
Shared configuration in the repo
- Commit the agent config: AGENTS.md (one file every tool can read, with a CLAUDE.md that imports it if needed), shared skills, subagents, hooks,
.mcp.json, and permission allow/deny lists in project settings. Personal preferences go in gitignored local files (CLAUDE.local.md,settings.local.json). Anthropic docs 2026 - Org-level policy where available: Claude Code supports a managed CLAUDE.md and managed settings (for example
permissions.deny, enforced sandboxing) that users can't override. Anthropic docs 2026 Cursor has Team Rules. Cursor docs 2026 - Own the config like code: review changes to AGENTS.md and hooks in PRs, prune regularly, and add a line when the agent repeats a mistake.
Conventions for AI-authored PRs
- Accountability: a human author owns every PR, whoever typed it. "The AI wrote it" is never an answer in review or in an incident.
- Disclosure: mark agent-authored or agent-heavy PRs (label, trailer, or the tool's default attribution) so reviewers calibrate, and so you can measure later.
- Evidence in the description: what/why, how it was verified (commands and output, screenshots), risks, and assumptions made.
- Size and scope limits: one concern per PR. Reject unrelated refactors and drive-by reformatting.
- No self-approval loops: an agent's PR can't be approved by an agent alone. Keep required human review and branch protection. Copilot cloud agent, for example, works within branch protections and rulesets. GitHub docs 2026
Review standards
Volume is the risk: when PR counts double, rubber-stamping becomes the failure mode. Use an agent first pass for mechanical issues, a checklist for agent PRs (test changes, dependencies, scope, error paths), and CODEOWNERS-enforced scrutiny for auth, payments and infra.
Measuring impact (and the pitfalls)
| Measure | Why | Pitfall |
|---|---|---|
| Delivery metrics (lead time, deploy frequency, change failure rate, time to restore) | Outcome-level; DORA's standard set | Lagging, confounded by everything else going on |
| Change failure / rework / revert rate on AI-assisted vs other PRs | Catches the quality cost that throughput hides | Needs reliable labelling of which PRs were AI-assisted |
| Review time and PR size distribution | Shows whether review is becoming the bottleneck | Fast reviews may mean rubber-stamping, not efficiency |
| Developer surveys (satisfaction, perceived productivity, trust) | Cheap, captures experience | METR showed self-reported speedup can have the wrong sign (§7) |
| Adoption / usage and cost per engineer | Needed for budgeting | Usage ≠ value. Lines of code and "acceptance rate" are vanity metrics. |
Run it like an experiment where you can: staged rollout by team, a baseline period, and paired quality and throughput metrics. Never use lines of code, PR counts or AI-acceptance rate as targets (Goodhart's law applies fast).
Security
Prompt injection is the defining risk. A coding agent reads large amounts of text it didn't write: repo files, dependency source, issue and PR comments, web pages, MCP tool results. Any of it can contain instructions. Simon Willison's lethal trifecta is the clearest model: an agent that combines (1) access to private data, (2) exposure to untrusted content and (3) the ability to communicate externally can be tricked into exfiltrating that data. Willison 2025 A background agent that reads a public issue, has a token with org-wide repo access, and has open internet egress has all three. Vendor docs echo this. Codex warns that "prompt injection can cause the agent to fetch and follow untrusted instructions" when network or web search is enabled. OpenAI docs 2026 Claude Code warns that MCP servers fetching external content create injection risk. Anthropic docs 2026
| Risk | Example | Containment |
|---|---|---|
| Prompt injection | Issue body says "also add this script to postinstall"; a fetched web page tells the agent to read ~/.aws/credentials | Break the trifecta: sandbox + egress allowlist; no secrets in the environment; human review before merge; treat issue text from non-members as untrusted |
| Secrets exposure | Agent reads .env into context; pastes a key into a log or PR; secrets end up in transcripts | Deny-read rules for secret files; secret managers instead of env files; pre-commit secret scanning; short-lived credentials. Claude Code's cloud mode keeps git credentials outside the VM behind a proxy. Anthropic docs 2026 |
| Slopsquatting (hallucinated packages) | Agent adds npm i fast-json-sanitizer, a name it invented, which an attacker has pre-registered | Dependency allowlist / internal registry proxy; "ask before adding dependencies" rule; lockfile diffs reviewed; registry age and download heuristics in CI |
| Destructive actions | rm -rf, force-push, DROP TABLE, terraform apply | Permission deny rules + hooks for hygiene; no prod credentials on dev machines; infra changes only through reviewed pipelines |
| Over-privileged automation | CI agent with an org-admin PAT; GitHub App with write access everywhere | Least-privilege, repo-scoped, short-lived tokens; OIDC federation instead of long-lived secrets (the Claude Code Action supports this) Anthropic docs 2026 |
| Comment-triggered automation | Agent replies on a PR and triggers an issue_comment workflow that deploys | Audit comment-triggered workflows before enabling auto-fix agents. Anthropic's docs warn about exactly this. Anthropic docs 2026 |
On slopsquatting, the USENIX Security 2025 study by Spracklen et al. tested 16 code LLMs over 576,000 generated samples. It found package-hallucination rates of at least 5.2% for commercial models and 21.7% for open-source models, with 205,474 unique invented package names. Spracklen+ 2025 Those rates were measured on 2024-era models in a specific prompting setup, so treat them as evidence that the risk is real, not as current rates.
Assuming there's no risk until merge. The agent runs code (install scripts, tests) during the task, so exfiltration happens at execution time. Sandbox execution, not just the merge.
Cost management
Drivers: context per turn, model tier, reasoning effort, parallel agents, runaway loops. Controls: budgets and alerts, cheaper models for subagents and fan-out, --max-turns and timeouts in CI, concise memory files, usage dashboards. CI runs consume both Actions minutes and tokens. Anthropic docs 2026 Anthropic docs 2026 Concrete per-seat prices change often, so look them up rather than quoting from memory.
"Roll out agents to a 50-person team?" Pilot with baseline metrics, then a paved road (shared config, sandbox defaults), a security review, PR conventions, enablement, paired throughput and quality measurement, and expansion team by team. Full model answer in the question bank. Flag that review capacity and CI speed become the new bottlenecks.
- The lethal trifecta for AI agents: the threat model for injection plus exfiltration.
- We Have a Package for You! (USENIX Security 2025): the package-hallucination / slopsquatting study.
- Claude Code security: a vendor's description of isolation and data handling.
- Codex approvals and security: sandbox modes and network defaults.
7. What the productivity evidence actually says
Expect to be asked "does AI actually make engineers faster?" The defensible answer is "it depends on the task, the developer, the codebase and the year, and the best evidence is mixed". Then show you know the studies.
| Study | Design | Headline | Key caveats |
|---|---|---|---|
| Peng et al. (GitHub/Microsoft), 2023 arXiv 2023 | Controlled experiment: implement an HTTP server in JavaScript, with vs without Copilot (autocomplete era) | Treatment group finished ~55.8% faster | One small greenfield task; a lab setting; authors affiliated with the vendor; pre-agent tooling |
| Cui, Demirer, Jaffe, Musolff, Peng, Salz (three field RCTs) MIT working paper | Randomized Copilot access at Microsoft, Accenture and a Fortune 100 company; 4,867 developers | ~26% more completed tasks (SE ≈ 10 pts); larger gains and adoption among less-experienced developers | Outcome proxies are PRs, commits and builds, not value or quality; autocomplete-era tool; imperfect compliance |
| Google enterprise RCT arXiv 2024 | 96 Google engineers, an enterprise-grade task, internal AI tooling, summer 2024 | ~21% less time, with a wide confidence interval | A single lab task; authors caution the effect "will not necessarily apply more broadly" |
| METR, early-2025 METR 2025 arXiv 2025 | RCT: 16 experienced maintainers of large mature open-source repos, 246 real issues randomly assigned AI-allowed or not; mainly Cursor Pro with Claude 3.5/3.7 Sonnet | 19% slower with AI (CI roughly +2% to +39%). Developers predicted 24% faster and afterwards believed they'd been 20% faster | Small n; experts in repos they know deeply (a setting where AI helps least); early-2025 tools; METR explicitly says it does not show AI fails to help most developers |
| METR follow-up (Feb 2026) METR 2026 | 57 developers (10 returning, 47 new), 143 repos, 800+ tasks, late-2025 tools | Point estimates now lean toward speedup: about −18% time for returning devs (CI −38% to +9%) and −4% for new recruits (CI −15% to +9%) | Both CIs include zero. Strong selection effects: developers declined to participate or withheld tasks they didn't want to do without AI, which biases the estimate. METR says the data "gives us an unreliable signal" and is changing its study design. |
| DORA 2025: State of AI-assisted Software Development Google Cloud 2025 DORA 2025 | Survey of nearly 5,000 professionals plus 100+ hours of qualitative data | ~90% use AI at work; 80%+ report higher productivity; ~30% have little or no trust in AI code. AI adoption now correlates positively with throughput but still negatively with delivery stability. AI acts as an "amplifier". | Correlational, self-reported; describes associations, not causal effects; the capabilities model is DORA's framework |
Sort the studies by how much context the human already has. Greenfield tasks with less-experienced developers show big gains. Experts in codebases they know intimately, on issues full of tacit requirements, show little gain or a loss, because time saved typing goes into prompting, waiting and reviewing.
Citing vendor headline numbers ("55% faster!") as general truth, or citing METR's 19% as proof that "AI makes you slower". Both are over-generalizations of narrow studies. The METR perception gap is the most transferable finding: self-reported productivity is unreliable, so measure outcomes.
New studies appear every few months, and 2026 results with late-2025 agentic tools may already have superseded some of this table. Before an interview, search for the latest METR, DORA and peer-reviewed RCT results, and quote figures only with their caveats.
How to talk about it: "The controlled evidence ranges from big speedups on greenfield lab tasks to a measured slowdown for experts in mature repos, and METR's 2026 follow-up says selection effects now make that design unreliable. DORA's survey work suggests AI amplifies existing engineering practices: throughput up, stability not necessarily. So I treat it as a capability whose value depends on workflow. I measure outcomes rather than trusting how fast I feel, and I invest in what makes agents effective: tests, specs, fast CI and clear conventions." This shows you've read the research and won't overclaim.
- METR early-2025 study: read the caveats section, not just the headline.
- METR Feb 2026 update: a good discussion of why measuring AI uplift is now hard.
- DORA 2025 report: the amplifier framing and the AI capabilities model.
- Google enterprise RCT: a careful lab study in an industrial setting.
8. Vibe coding vs agentic engineering
Vibe coding was coined by Andrej Karpathy in February 2025 for a style where you describe what you want, accept the AI's code without really reading it, and steer by running it and prompting again. Collins named it Word of the Year for 2025. Wikipedia 2026 Karpathy 2025 Simon Willison proposed vibe engineering (October 2025) for the opposite end: experienced engineers using agents to go faster while remaining accountable for the software, which leans heavily on tests, planning, documentation, version control, CI and code review. In a February 2026 update he noted that "agentic engineering" seemed to be becoming the more popular term. Willison 2025
| Vibe coding | Agentic engineering | |
|---|---|---|
| Reads the code? | Mostly no; judges by behaviour | Yes. Reviews every diff that ships |
| Verification | "It seems to work when I click around" | Tests, typecheck, CI, reviewer, evidence |
| Specs | Conversational, emergent | Written specs, acceptance criteria, non-goals |
| Who's accountable | Nobody, really | The engineer, fully |
| Good for | Prototypes, demos, personal tools, learning, exploring an idea, throwaway scripts | Production systems, shared codebases, anything with users, data or money |
| Main risk | Security holes, unmaintainable code, prototype becomes production | Over-process on trivial tasks; review fatigue |
Most engineers do both, on purpose: vibe-code a prototype to test an idea, then rebuild it properly. The skill is noticing when a vibe-coded artifact is about to become production.
"What do you think about vibe coding?" Don't dismiss it and don't celebrate it. Say it's a legitimate tool for exploration and throwaway work and a liability for production, and that what matters is accountability: "if it ships under my name, I've read it, it's tested, and I can explain it." Then describe your own boundary, with an example of each.
- Vibe engineering (Willison): the practices that separate accountable agent use from vibes.
- Vibe coding (Wikipedia): origin, definition and the debate around it.
Interview question bank
Walk me through how you use AI coding tools day to day.
Three modes: autocomplete while typing, an interactive terminal agent for feature work, debugging and codebase Q&A, and a background agent for well-scoped tickets I review asynchronously. For a feature, I explore in plan mode, have the agent interview me into a short spec with acceptance criteria and non-goals, then start a fresh session that writes failing tests first and implements until they pass, running typecheck and lint itself. A fresh-context reviewer checks the diff against the spec. Then I review it, tests first. Independent tasks run in parallel worktrees. Novel design and security-critical code I keep close. One real catch: an agent "fixed" a flaky test by adding a retry. Reading the test diff first exposed it, and we found the actual race.
How do you make sure agent-written code is correct?
Defense in depth. Before: a spec with acceptance criteria, and tests committed before implementation so any change to them is visible. During: the agent runs the real checks and pastes the evidence. For unattended runs, a Stop hook enforces it. After: an independent reviewer in a fresh context, then my review, starting with test changes (loosened assertions, skips, special-casing), then error paths and scope. System: CI, branch protection, required human approval. And I own it: if I can't explain a line, it doesn't ship.
How would you roll out coding agents to a 50-person engineering team?
Pilot first: 5–8 volunteers of mixed seniority on real work for 4–6 weeks, with a baseline captured before. Meanwhile build a paved road: a template AGENTS.md, shared skills and hooks, sandbox and permission defaults, approved MCP servers, all committed and reviewed like code. Run a security review (data retention, token scopes, egress, injection threat model, dependency policy). Set conventions: disclose AI-assisted PRs, include verification evidence, cap PR size, require human approval. Enable people with demos and a channel for sharing skills. Measure paired throughput and quality (lead time, change failure and revert rate, review time) plus cost, and distrust self-reports. Expand team by team, and expect review capacity and CI speed to become the bottlenecks.
When do you NOT use an agent?
When explaining the task takes longer than doing it. When I can't verify the result cheaply: subtle concurrency, numerics without reference outputs, security-critical auth. There I may use the agent to explore or draft, but I write and reason through the final code. When the work is mostly deciding rather than typing, though agents are useful sparring partners for that. When the task would touch production credentials and can't be sandboxed. And when I'm deliberately learning something, because writing it is how I learn.
How do you write a good CLAUDE.md / AGENTS.md?
Short, specific and checkable. Include what the agent can't infer: exact build, test and lint commands (including how to run one test), conventions that differ from defaults, architectural rules, environment gotchas, frozen areas, PR etiquette. Leave out what it can read from the code, generic advice, long docs (link them) and anything that changes often. The test for every line: would removing it cause mistakes? Keep it to a couple of hundred lines at most. Procedures go into skills, path-specific rules into scoped rule files, must-always rules into hooks. Use AGENTS.md as the cross-tool file and import it from CLAUDE.md. Review and prune it like code, and add a line when the agent repeats a mistake.
How do you manage context in a long session?
Performance degrades as the window fills, so: one task per session and clear between tasks. Push file-heavy research into subagents. Keep tool output small (filtered test runs, log tails). Keep durable state in files like PLAN.md so compaction or a fresh session loses nothing. Compact with a focus instruction, and put "what to preserve" in the memory file. After two failed corrections on the same issue, clear and re-prompt with what I learned, because failed attempts bias the model. When behaviour seems off, check what's loaded. Bloated memory files and many MCP tools eat budget silently.
What are the security risks of coding agents and how do you contain them?
The central one is prompt injection: the agent reads untrusted text (issues, dependency code, web pages, MCP results). Combined with private data and an outbound channel, that's the lethal trifecta. Others: secrets in context or logs, destructive commands, over-privileged CI tokens, slopsquatting, and comment-triggered workflows. Containment: a sandbox with filesystem and network isolation plus an egress allowlist; no long-lived secrets in the environment (proxy-brokered, short-lived, repo-scoped tokens, OIDC in CI); deny-read rules for secret files; a dependency allowlist and lockfile review; hooks for hygiene, while remembering they're not a boundary; and human review before merge. Code executes during the task, so isolate execution, not just merges.
Explain skills vs subagents vs hooks vs MCP. When would you use each?
Skill: a SKILL.md folder whose name and description load at startup, while the body and scripts load on demand. Use it for procedures needed sometimes. It's an open standard many tools read. Subagent: an agent with its own context, tools and model that returns a summary. Use it for context-heavy research, independent review or fan-out, not for coupled writes. Hook: a deterministic script at a lifecycle event. Use it for anything that must always happen (format, lint, block force-push, don't stop until typecheck passes). MCP: external tools and data. Use it when a CLI won't do, and watch context cost and injection risk. In short: knowledge in memory files and skills, enforcement in hooks and permissions, reach through MCP or a CLI.
How does a coding agent actually work under the hood?
An LLM in a tool-use loop. The harness sends the system prompt, memory files, skill descriptions, tool schemas and the conversation. The model emits tool calls (read, grep/glob, edit, shell, web). The harness checks permissions, runs them in a sandbox, appends results, and repeats until the model stops or a budget or interrupt ends it. Edits are usually exact-match replacements validated by the harness, with snapshots for undo. Context comes mostly from agentic search, often in a read-only subagent. When the window fills, old tool outputs are dropped and the conversation is summarized. Verification runs the project's own checks. The harness matters as much as the model.
An agent tells you "all tests pass." What do you do?
Ask for the command and its output, ideally already required by the prompt or a Stop hook. Then check the ways the claim can be true and still wrong. Did it run the whole relevant suite or one file? Did it modify, skip or loosen tests, or regenerate snapshots? Do the new tests fail if I revert the implementation? Is there special-casing of test inputs? For anything important, I re-run the check or rely on CI. A reviewer subagent with this checklist catches most of it cheaply.
How do you run multiple agents in parallel without chaos?
Isolation and independence. Each agent gets its own worktree, clone or cloud sandbox, with its own env files, dependencies and ports. I parallelize only tasks that are independent and come with their own checks. Coupled changes stay in one session, because parallel writers make inconsistent implicit decisions. Specs written up front mean agents rarely need me mid-task, and I batch reviews. The real limit is my review bandwidth: for me, 2–4 concurrent agents. A second session can also act as an unbiased reviewer of the first.
How do you use agents for a large refactor or migration?
If it's purely syntactic, use a deterministic codemod (ast-grep, jscodeshift, OpenRewrite), possibly written by the agent. If it needs judgment: have the agent produce a work list and a precise per-item prompt with a verification step. Run headless on 2–3 items, refine from the failures, then fan out with tightly scoped tool permissions, in parallel worktrees if supported. Each item reports OK or FAIL. Then run the full suite and typecheck, spot-check a random sample of diffs, handle failures by hand, and land the change in reviewable chunks.
What does the research say about AI coding productivity?
Mixed, and dependent on context. The 2023 Copilot lab study found about 56% faster completion of a small greenfield task. Field RCTs across three companies found about 26% more completed tasks, with bigger gains for juniors. Google's internal RCT found about 21% less time, with wide uncertainty. METR's 2025 RCT with expert open-source maintainers found a 19% slowdown, while participants believed they'd been 20% faster. Its 2026 follow-up leans toward speedup, but the CIs span zero and METR calls the signal unreliable because of selection effects. DORA 2025 sees AI as an amplifier: throughput up, stability not. So value depends on task, familiarity and workflow, and self-reports can't be trusted. Measure outcomes.
What's the difference between vibe coding and how you work?
Vibe coding (Karpathy's term) means accepting output without reading it and steering by behaviour. I use it deliberately for spikes, demos and throwaway tools. For production I do what Willison calls vibe engineering, now usually called agentic engineering: specs, tests first, the agent proving its work, and me reviewing every diff, because I'm accountable for what ships. The skill is choosing the mode on purpose and noticing when a prototype is about to become production.
How do you review a PR opened by a background agent?
With more suspicion than a colleague's PR, because it reads fluently whether or not it's right. Task and spec first, then test changes (do the new tests fail without the change, were old ones weakened?), then the core diff, then ripple effects. I check for scope creep, new dependencies (real? needed? reputable?), error handling and silent assumptions. I look at the attached evidence and re-run anything suspicious. If the approach is wrong, I close the PR and fix the spec rather than steer through comments. Usually the task wasn't specified well enough to delegate.
How do you keep costs under control?
The drivers are context size per turn, model tier and reasoning effort, parallel agents, and retries or loops. Keep memory files concise and MCP tools few, since they're paid for every turn. Use cheaper models for subagents, search and mechanical fan-out, and strong models for planning and hard debugging. Set turn limits and timeouts in CI, clear context instead of dragging long sessions along, and set budgets and dashboards per team. Track cost per merged, non-reverted change. A stronger model that gets it right first time is often cheaper overall.