Model Evaluation & Interpretability
Two ways to answer "what does this model actually do?" Evaluation looks from the outside: what you measure, how you score it, and how much to trust the number. Interpretability looks from the inside: which internal representations and circuits produce the behaviour. This page treats evaluation the way a model builder sees it (pretraining and post-training checkpoints, public benchmarks, leaderboards, capability and safety evals). Product evals for LLM applications (task rubrics, golden sets, online metrics) are covered separately in B5. Interviewers use this area to tell apart people who read leaderboards from people who can say why a 2-point gain on a 500-question benchmark may be noise, why a judge model prefers longer answers, or what a sparse autoencoder is optimizing.
TL;DR: the 8–12 things to be able to say out loud
- Perplexity is \(\exp\) of the mean per-token negative log-likelihood. It is great for comparing checkpoints that share a tokenizer and data distribution, but it doesn't compare across tokenizers (use bits-per-byte) and it predicts downstream task quality only loosely, especially after RLHF.
- Benchmarks follow a lifecycle: introduced hard, then optimized against, contaminated and saturated within roughly 1–3 years. MMLU, GSM8K and HumanEval are saturated. SWE-bench Verified was retired by OpenAI in early 2026 over flawed tests and contamination. Frontier signal now comes from HLE, FrontierMath, ARC-AGI-2/3, Terminal-Bench 2.0, SWE-bench Pro and agentic or long-horizon evals.
- pass@k should be computed with the unbiased estimator \(1-\binom{n-c}{k}/\binom{n}{k}\) from \(n \ge k\) samples, not by drawing \(k\) samples once. maj@k (self-consistency) is a different thing: it's a single answer chosen by voting.
- Arena leaderboards fit a Bradley–Terry model to pairwise votes: \(P(A \succ B)=\sigma(\beta_A-\beta_B)\). A 100-point Elo gap means about a 64% win rate. Arenas are biased by style, by who gets sampled, and by selective disclosure.
- The scoring protocol changes the number: zero- vs few-shot, prompt format (up to tens of points on older models), log-likelihood vs generative scoring, and answer extraction. Always report the protocol.
- Contamination: test items leak into pretraining data. You detect it with n-gram overlap, membership inference (Min-K% Prob), exchangeability tests, or by building fresh mirror sets (GSM1k) and live benchmarks (LiveCodeBench, LiveBench).
- LLM-as-judge has position, verbosity and self-preference biases. Mitigations: swap order and average, control for length, use a judge from a different family or a jury of judges, use reference answers, and calibrate against human labels.
- Statistics: the standard error of accuracy is \(\sqrt{p(1-p)/n}\). With n=500 and p=0.5 that is ±4.4 points at 95% confidence. With n=30 (one year of AIME) it is about ±18 points. Use paired tests and clustered SEs, and resample when the model is stochastic.
- Reasoning models should be compared on accuracy-vs-tokens (or cost) curves at matched budgets, not on single numbers.
- Interpretability toolkit: linear probes (correlational), logit lens and tuned lens, attention patterns (suggestive, not explanations), and activation patching (causal). Mechanistic interp views the model as a residual stream that components read from and write to.
- Superposition: models represent more features than they have dimensions, which is why neurons are polysemantic. Sparse autoencoders learn an overcomplete sparse dictionary, with loss = reconstruction + λ·sparsity. Transcoders and cross-layer transcoders let you build attribution graphs of computation.
- Steering vectors add a direction (for example a mean activation difference) to the residual stream to change behaviour. ROME/MEMIT edit facts through rank-one updates to MLP weights. CoT is often not faithful: reasoning models verbalize the hints they actually used less than 20% of the time in many settings.
Part 1 · Model evaluation: why it's harder than it looks
Evaluating a classifier on ImageNet was a solved social problem: one fixed test set, one metric, one protocol. LLM evaluation breaks each of those assumptions:
- Open-ended outputs. A model can be correct in many surface forms, so scoring needs answer extraction, execution (code), or a judge. Each of these adds error.
- Many protocol degrees of freedom. System prompt, few-shot examples, chat template, sampling temperature, max tokens, reasoning effort, tool access and scaffolding all change results. A "score" really belongs to a (model, protocol) pair.
- Training data is the internet. Public test sets end up in pretraining corpora (contamination), and labs tune hyperparameters, data mixes and RL environments against the benchmarks they report (Goodhart).
- Capabilities emerge and shift. Benchmarks saturate fast. The interesting capabilities move toward long-horizon agentic tasks that are expensive to run and noisy to grade.
- Small, correlated test sets. Many headline benchmarks have a few hundred items (GPQA Diamond about 198, AIME 30 per year, SWE-bench Verified 500), so naive comparisons are swamped by noise.
A model builder runs evals at several points, each with its own goal:
| Stage | What you measure | Typical tools | Main failure mode |
|---|---|---|---|
| Pretraining (during the run) | Held-out loss/perplexity, cheap log-likelihood multiple-choice (HellaSwag, ARC-Easy, MMLU cloze), small generative math/code sets | Fast in-loop evals every N steps, scaling-law fits | Metrics that are flat until a threshold ("emergence"), which leaves no signal for small-scale ablations |
| Mid/post-training | Instruction following, chat quality, math/code (generative), tool use, refusal behaviour | Generative harnesses, LLM judges, internal preference evals | Overfitting to the judge or to public benchmarks, prompt-format sensitivity |
| Release decision | Frontier capabilities, agentic tasks, human preference, dangerous capabilities | Agent harnesses (sandboxes), arenas, red teams, third-party evaluators | Elicitation gaps (under-estimating what the model can do), cost |
| Post-release | Real usage, regressions, arena rankings | Arena-style platforms, telemetry (that's B5 territory) | Selection bias, style effects |
"How would you evaluate a new base model checkpoint vs. the previous one?" A strong answer separates cheap, low-variance, in-loop signals (held-out loss on several domains, bits-per-byte, log-likelihood multiple-choice) from expensive generative evals. It holds the protocol fixed across checkpoints, reports confidence intervals and paired differences, checks contamination of any new data, and picks metrics that are continuous at the current scale (log-prob of the correct answer rather than exact match) so small ablations still show signal.
Perplexity and its limits
For a held-out token sequence \(x_1,\dots,x_N\), the average negative log-likelihood (cross-entropy, in nats) is
$$\mathcal{L} = -\frac{1}{N}\sum_{i=1}^{N}\log p_\theta(x_i \mid x_{\lt i}), \qquad \mathrm{PPL} = e^{\mathcal{L}}.$$Intuitively, perplexity is the effective branching factor: PPL = 10 means the model is, on average, as uncertain as if choosing uniformly among 10 tokens. It's the pretraining objective itself, so it's smooth, cheap, low-variance, and it is what scaling laws are fit to (see the scaling-laws page).
Why you can't stop at perplexity
- Tokenizer dependence. PPL is per token. A model with a larger vocabulary predicts fewer, "bigger" tokens, so per-token PPL isn't comparable across tokenizers. Normalize to bits per byte (or per character): \(\mathrm{BPB} = \frac{N_\text{tokens}}{N_\text{bytes}}\cdot\frac{\mathcal{L}}{\ln 2}\). This is what tokenizer-agnostic comparisons use.
- Distribution dependence. PPL on web text says little about PPL on code or math. Report it per domain. A data-mix change can improve one domain and hurt another while the average barely moves.
- Weak link to task quality. Most of the likelihood mass is on "easy" tokens (function words, boilerplate). A model can gain on PPL by getting better at predictable tokens while getting no better at the rare, decisive tokens of a reasoning chain. Downstream accuracy often looks like a sharp function of loss, and much of that sharpness comes from discontinuous metrics such as exact match. Schaeffer et al. showed that many "emergent abilities" turn smooth when you switch to continuous metrics like token edit distance or log-prob of the correct answer. Schaeffer+ 2023
- Post-training breaks it. RLHF and RL on verifiable rewards deliberately move the model away from the pretraining distribution. Chat models become low-entropy and are often worse calibrated, so their PPL on generic text gets worse even as they become far more useful. Perplexity is a pretraining metric.
- Context-length artefacts. PPL on long documents mostly reflects local prediction. A model can have excellent long-context PPL and still fail to retrieve a fact from 100k tokens back, which is why RULER-style tasks exist.
- Contamination inflates it silently. Memorized documents have very low PPL. Turned around, this becomes a contamination detector (Min-K% Prob, below).
Loss is the right training signal because it's dense: every token gives gradient. It's an imperfect evaluation signal because the tasks you care about depend on a tiny number of high-stakes tokens (the final answer, the right API call). Benchmarks exist to put weight on those tokens.
Comparing Llama-family and GPT-family perplexities directly, or comparing PPL between a base and a chat model and concluding that the chat model "got worse". Normalize by bytes and only compare like with like.
- Are Emergent Abilities of LLMs a Mirage?: how metric choice creates or removes apparent discontinuities.
- Lessons from the Trenches on Reproducible Evaluation (Biderman+ 2024): log-likelihood vs generative scoring, normalization choices, from the lm-eval-harness maintainers.
The benchmark landscape (as of late 2026)
The table groups widely cited benchmarks by what they claim to measure. "Status" is a rough judgement of whether the benchmark still separates frontier models. Frontier scores move monthly, so treat the numbers as indicative.
Benchmark status changes fast. The saturation notes reflect public reporting up to about September 2026. Leaderboard trackers often disagree by many points because they use different protocols (tools vs no tools, text-only subsets, reasoning effort). Check the official leaderboard before quoting a number.
| Category | Benchmark | What it is | Scoring | Status |
|---|---|---|---|---|
| Knowledge | MMLU Hendrycks+ 2020 | About 14k four-way multiple-choice questions over 57 subjects | Accuracy (log-likelihood or generative) | Saturated: frontier models above 90%. About 6.5% of questions have errors, per MMLU-Redux Gema+ 2024 |
| MMLU-Pro Wang+ 2024 | Harder, reasoning-heavier, 10 options (so random guessing gives 10%) | Accuracy, usually with CoT | Close to saturation at the frontier | |
| Math | GSM8K Cobbe+ 2021 | About 8.5k grade-school word problems (1.3k test) | Exact match on the final number | Saturated. Useful only for small models and contamination studies |
| MATH Hendrycks+ 2021 | 12.5k competition problems, 5 difficulty levels. MATH-500 is the common subset | Exact match after normalizing the answer | Saturated at the frontier (above 95% on MATH-500 for reasoning models) | |
| AIME (yearly) / FrontierMath Glazer+ 2024 | AIME: 30 integer-answer olympiad-qualifier problems per year. FrontierMath: unpublished research-level problems | Exact match. Often avg@k or pass@1 averaged over samples | AIME 2024/2025 near saturation for top reasoning models. Each year's fresh set is used for contamination control. FrontierMath's top tiers still discriminate | |
| Science | GPQA (Diamond) Rein+ 2023 | "Google-proof" graduate-level biology, physics and chemistry multiple-choice. Diamond subset of about 198 questions | Accuracy | Approaching its ceiling: frontier models above the expert-validator baseline of about 65–70%. The small n makes it noisy |
| Code | HumanEval Chen+ 2021 | 164 hand-written Python function problems with unit tests | pass@k | Saturated and widely contaminated |
| MBPP Austin+ 2021 | About 1k crowd-sourced entry-level Python problems | pass@k | Saturated | |
| LiveCodeBench Jain+ 2024 | Problems pulled continuously from LeetCode, AtCoder and Codeforces, tagged by release date | pass@1. Evaluate only on problems released after the model's cutoff | Active. The date-windowing is a contamination control | |
| SWE-bench / Verified Jimenez+ 2023 | Real GitHub issues from 12 Python repos. The agent edits the repo, and hidden tests decide pass/fail. Verified is a 500-task subset that humans screened | % resolved | OpenAI stopped reporting Verified in Feb 2026, citing flawed tests and evidence that frontier models had memorized gold patches OpenAI 2026 | |
| SWE-bench Pro Deng+ 2025 | 1,865 longer-horizon tasks from 41 repos, with public, held-out and commercial (private) splits | % resolved | Active. Now recommended by OpenAI in place of Verified | |
| Agents | τ-bench / τ²-bench Yao+ 2024 Barres+ 2025 | A tool-using agent serves a simulated user (retail, airline, telecom) under domain policies. τ² adds dual control, where the user also acts on the environment | Task success checked against the final database state. pass^k = all k trials succeed (a reliability measure) | Active. pass^k exposes inconsistency that pass@1 hides |
| Terminal-Bench 2.0 Merrill+ 2026 | 89 hard tasks in containerized terminals (sysadmin, builds, crypto, scientific porting) with human-written solutions and tests | % tasks passed | Active. The paper reports frontier agents below 65% at publication | |
| OSWorld Xie+ 2024 | 369 real desktop tasks (Ubuntu/Windows apps) for multimodal computer-use agents | Execution-based checkers | Active, rising fast. Human baseline about 72% | |
| BrowseComp Wei+ 2025 | Hard-to-find facts that need persistent web browsing. Answers are short and easy to verify | Accuracy | Active | |
| METR time horizon Kwa+ 2025 | A meta-metric: the length of task, in human-expert time, that the AI completes with 50% success | Logistic fit of success against log(human time) | Widely cited. Reported doubling about every 7 months since 2019 | |
| Long context | Needle-in-a-haystack Kamradt 2023 | Plant one fact at depth d in a context of length L, then ask for it. Results are shown as a heatmap | Retrieval accuracy | Saturated. Literal retrieval is too easy and says little about real long-context use |
| RULER Hsieh+ 2024 | Synthetic tasks with configurable length: multi-needle, multi-hop tracing, aggregation, QA | Accuracy versus length. Defines the "effective context length" | Useful. It shows most models degrade well before their advertised window | |
| Frontier / hard | Humanity's Last Exam Phan+ 2025 | 2,500 expert-written closed-ended questions, some multimodal, filtered so that frontier models failed them when the set was built | Accuracy plus calibration error | Frontier models went from single digits (Jan 2025) to roughly 40–65% by late 2026, depending on tracker and tool use |
| ARC-AGI-1/2 Chollet 2019 Chollet+ 2025 | Grid-transformation puzzles that test few-shot abstraction from "core knowledge" priors | Exact-match grids, often with an efficiency (cost per task) axis | ARC-AGI-1 largely beaten. ARC-AGI-2's Kaggle 2025 top score was 24% on the private set Chollet+ 2026 | |
| ARC-AGI-3 ARC Prize 2026 | Interactive turn-based game environments where the agent must explore, infer goals and plan, with no instructions | Action-efficiency relative to humans | Launched March 2026. At release, frontier systems scored very low while humans solved all environments. Progress is fast and contested | |
| Human preference | Chatbot Arena / LMArena Chiang+ 2024 | Crowd users chat with two anonymous models and vote | Bradley–Terry "Elo" with CIs, plus style-controlled variants | Very influential. Bias and gaming critiques in Singh+ 2025 |
| Factuality | SimpleQA Wei+ 2024 | About 4.3k short fact-seeking questions written adversarially against GPT-4 | Correct / incorrect / not attempted, which rewards abstaining | Useful for measuring hallucination vs. abstention |
| Economic tasks | GDPval Patwardhan+ 2025, MLE-bench Chan+ 2024 | Real deliverables from professional occupations, graded by experts (GDPval). Kaggle competitions (MLE-bench) | Expert win rate vs. human professionals; medal rate | Newer, expensive to grade |
Notes on specific figures: the GPQA expert baseline and OSWorld human baseline are approximate values from the respective papers. HLE frontier ranges are taken from third-party trackers, which disagree with each other.
"Which benchmarks would you trust today and why?" Don't recite a list. Give criteria. (1) Headroom: frontier scores well below the ceiling, and the ceiling isn't capped by label noise. (2) Contamination resistance: held-out, private or continuously refreshed items, or an exact release-date cutoff. (3) Grading validity: execution or state checks rather than string match, and audited tests. (4) Enough items for the differences you care about. (5) Construct validity: does it measure the capability your product needs? Then name examples: SWE-bench Pro / Terminal-Bench for agentic coding, LiveCodeBench for algorithmic coding, HLE / FrontierMath for frontier reasoning, τ²-bench pass^k for reliability, and your own private evals above all.
- Epoch AI Benchmarking Hub: independently run results on hard benchmarks, with methodology notes.
- OpenAI: why SWE-bench Verified no longer measures frontier coding: a case study in benchmark death.
- Are We Done with MMLU?: how label errors cap the measurable ceiling.
Metrics: accuracy, pass@k, maj@k, pass^k, Elo
Accuracy and its variants
For multiple-choice or exact-answer tasks the score is just the mean of per-item 0/1 correctness. The details that matter: how you extract the answer (regex over the last "Answer: X", \boxed{} parsing, symbolic equivalence via sympy for math), how you handle refusals or unparseable outputs (count them as wrong, and report the parse-failure rate), and whether you report pass@1 averaged over several samples (often written avg@k, mean@k or pass@1 with n samples). For small sets like AIME, averaging over 8–64 samples per problem reduces sampling variance. It does not reduce the variance that comes from having only 30 problems.
pass@k and the unbiased estimator
pass@k is the probability that at least one of k independent samples is correct. It models a setting where a verifier, such as unit tests, can pick the good sample. The naive approach (draw exactly k samples, check if any pass) has high variance. Codex instead draws \(n \ge k\) samples per problem, counts the \(c\) correct ones, and uses the unbiased estimator Chen+ 2021:
$$\widehat{\text{pass@}k} \;=\; \mathbb{E}_{\text{problems}}\left[\,1-\frac{\binom{n-c}{k}}{\binom{n}{k}}\,\right].$$Read it as "1 minus the probability that a random size-k subset of the n samples contains no correct one". Computing the binomials directly overflows, so use the stable product form \(1-\prod_{i=n-c+1}^{n}\left(1-\frac{k}{i}\right)\) (and return 1 if \(n-c \lt k\)). Codex used n = 200 for k up to 100.
Why not just compute \(1-(1-\hat p)^k\) with \(\hat p = c/n\)? Because it's biased: \(1-(1-p)^k\) is concave in \(p\), so plugging in a noisy \(\hat p\) systematically underestimates the true value (Jensen's inequality). The combinatorial form is exactly unbiased for \(k \le n\).
Practical notes on pass@k: temperature matters (higher T helps large-k, hurts k=1, so the Codex paper tuned T per k); pass@k with large k is only meaningful when a verifier exists at deployment; and it's increasingly used as a probe for "latent capability": if pass@256 is high but pass@1 is low, RL may be able to surface the capability. Whether RL with verifiable rewards expands pass@k at large k or only sharpens pass@1 is an active and contested research question as of 2025–2026.
pass^k (reliability)
τ-bench introduced pass^k: the probability that all k independent trials succeed. With per-task success rate \(p_t\) it is \(\mathbb{E}_t[p_t^{\,k}]\), estimated unbiasedly as \(\binom{c}{k}/\binom{n}{k}\) Yao+ 2024. A model with 60% pass@1 that's consistent per task (some tasks 100%, others 0%) keeps a high pass^8. One that succeeds on every task 60% of the time collapses to 0.6⁸ ≈ 1.7%. Customer-facing agents need pass^k, not pass@k.
maj@k (self-consistency)
maj@k samples k reasoning chains, extracts each final answer, and outputs the majority vote Wang+ 2022. Unlike pass@k it produces one answer with no oracle, so it's a fair deployable comparison. It needs answers that can be compared for equality (numbers, choices), and it plateaus once the most common wrong answer dominates. Variants weight the votes by a reward model or by confidence (best-of-N with an RM). When a lab reports "maj@64", it is spending 64× inference compute, so compare at matched compute.
| Metric | Question it answers | Needs oracle at deploy? | Typical use |
|---|---|---|---|
| pass@1 (avg over n) | Expected single-shot accuracy | No | Default headline number |
| pass@k | Can the model produce a correct answer at all within k tries? | Yes (tests, verifier) | Code with unit tests, latent-capability probes |
| pass^k | Does it succeed every time? | No | Agent reliability |
| maj@k / cons@k | Accuracy of the voted answer from k samples | No | Math and multiple-choice with test-time compute |
| best-of-N (RM) | Accuracy when a learned verifier picks one of N | Learned RM | Test-time scaling studies |
Elo and Bradley–Terry for arenas
Arena platforms collect pairwise votes. The Bradley–Terry (BT) model assigns each model a strength \(\beta_m\) and assumes
$$P(A \succ B) = \sigma(\beta_A-\beta_B) = \frac{1}{1+e^{-(\beta_A-\beta_B)}}.$$The strengths are fit by maximum likelihood, which is just logistic regression where each battle is a row with +1 for A's column and −1 for B's. They're reported on an Elo-like scale, \(R = 400\,\beta/\ln 10 + \text{const}\), so that \(P(A\succ B)=1/(1+10^{(R_B-R_A)/400})\). A 100-point gap is about a 64% win rate, and 200 points about 76%. Chatbot Arena moved from online Elo updates, which depend on the order of games, to batch BT fitting with bootstrap confidence intervals Chiang+ 2024. Ties are handled by counting half wins or with a Rao–Kupper/Davidson extension.
Known issues with arenas:
- Style confounds. Longer, markdown-heavy, more confident answers win votes. LMArena added "style control", which adds response length and markdown counts as covariates in the BT regression so that the β coefficients are style-adjusted.
- Selective disclosure and sampling asymmetries. Providers testing many private variants and publishing only the best one inflates scores through a winner's curse. Proprietary models also got more battles. Singh+ 2025 reported, for example, 27 private Meta variants tested before Llama 4.
- Prompt distribution. Votes reflect what arena users ask (lots of coding and casual chat), not your users.
- Non-transitivity. BT assumes a single "strength" dimension. Real preferences vary by category, hence category leaderboards.
"Model A has Elo 1300, B has 1280. Is A better?" The strong answer: a 20-point gap is about a 53% win rate. Check whether the bootstrap CIs overlap (with tens of thousands of votes they're often ±5–10 points, but much wider for new models). Ask whether the score is style-controlled, which category it's in, and whether it would hold on your traffic. Then run your own blind pairwise eval on in-domain prompts.
- Evaluating LLMs Trained on Code (Codex): the pass@k estimator derivation and numerically stable code.
- Chatbot Arena paper: BT fitting, CIs, and active sampling of pairs.
- The Leaderboard Illusion: how arena rankings get distorted.
Scoring protocols: shots, prompts, log-likelihood vs generation
Zero-shot vs few-shot
Few-shot prompting puts k solved examples in the context. For base models it's essential: it tells the model the format and the task, and 5-shot MMLU was the standard. For chat/instruct models, zero-shot with an instruction (often plus "think step by step" or native reasoning) is now the norm. Few-shot exemplars can even hurt reasoning models by constraining their reasoning format. Results are not comparable across shot settings, and the choice and order of exemplars add variance on their own.
Prompt sensitivity
Meaning-preserving formatting changes (separators, casing, spacing, "Answer:" vs "A:") swung LLaMA-2-13B few-shot accuracy by up to 76 points on some tasks, and the best format barely correlated across models Sclar+ 2023. Frontier instruction-tuned models are much more robust, but sensitivity of a few points is still common, which is the same size as many claimed improvements. GSM-Symbolic showed that changing names and numbers in GSM8K templates, or adding irrelevant clauses, reduces accuracy and increases variance, a sign of pattern-matching rather than robust reasoning Mirzadeh+ 2024. Good practice is to report the mean and spread over several prompt variants, and to use the model's own chat template.
Log-likelihood (cloze/MC) vs generative scoring
Log-likelihood scoring
Compute \(\log p(\text{choice}_j \mid \text{prompt})\) for each option and pick the argmax. You can score the full answer text ("cloze") or just the letter ("MC format").
- Cheap: one forward pass, no sampling, deterministic
- Works on small or base models that can't follow output formats yet, so it's ideal for pretraining ablations
- Needs length normalization (per token, per character, or relative to an unconditional prior), and the choice changes rankings
- Can't evaluate reasoning chains, and isn't possible for closed APIs without logprobs
Generative scoring
Sample (or greedy-decode) a full response, extract the answer, and compare.
- Matches how the model is actually used, and allows CoT and reasoning
- Brittle answer extraction: format failures get counted as knowledge failures
- Sampling adds variance, so use multiple samples
- Much more expensive, especially with long reasoning traces
The same model can rank differently under the two schemes. Small models often look much better under cloze log-likelihood, because the letter-MC format is itself a capability that emerges only at scale. The lm-eval-harness authors document these choices in detail Biderman+ 2024.
Comparing your number against a number in someone else's paper. Unless the harness version, prompt, shots, chat template, sampling settings and answer extraction all match, differences of several points are expected. Re-run the baselines yourself in the same harness.
- Quantifying LM sensitivity to prompt formatting (FormatSpread)
- GSM-Symbolic: templated perturbations as a robustness and contamination probe.
Contamination
How it happens
- Direct inclusion. Test sets sit on GitHub, Hugging Face or in papers, and web crawls ingest them, sometimes with answers.
- Indirect leakage. Blog posts, forum discussions, solution write-ups and translated or paraphrased versions. Exact-match filters miss these.
- Synthetic data echoes. A teacher model that saw the test set generates training data resembling it.
- Post-training and RL environments. Tuning on problem sets drawn from the same sources (competition archives, GitHub issues) as the benchmark. Repository-based benchmarks are especially exposed: SWE-bench tasks come from public repos with public fix commits.
- Adaptive overfitting without leakage. Choosing checkpoints, data mixes and hyperparameters because they score well on the benchmark. No test item is ever trained on, but the reported number is still optimistically biased.
Detection methods
| Method | How it works | Needs | Limits |
|---|---|---|---|
| N-gram / substring overlap | Search the training corpus for long n-gram matches with test items. GPT-4's report used 50-character substring matching OpenAI 2023, and Llama 3 reports similar token n-gram analyses Grattafiori+ 2024 | Training data access | Misses paraphrases and translations. Thresholds are arbitrary |
| Membership inference (Min-K% Prob) | Average log-prob of the k% least likely tokens in a text. Seen text has fewer very-surprising tokens Shi+ 2023 | Logprobs only | Noisy per example. Better for aggregate evidence |
| Exchangeability / order test | If the model assigns higher likelihood to the benchmark's canonical item order than to shuffled orders, it probably saw the file. This gives a statistical guarantee Oren+ 2023 | Logprobs | Only detects verbatim, ordered inclusion |
| Completion / regurgitation probes | Prompt with the first half of an item and check whether the model completes it verbatim (or reproduces the gold patch) | Black-box sampling | Absence isn't proof of cleanliness |
| Fresh mirror set | Build new items with the same distribution and difficulty, then compare. GSM1k found drops of up to 8% for some model families, with frontier models showing little overfitting Zhang+ 2024 | Annotation budget | Expensive, and hard to match difficulty exactly |
| Perturbation tests | Change numbers, names or option order, or rephrase. Accuracy that drops on the perturbed items suggests memorization | Templates | Perturbations can change difficulty |
| Time-split ("live") evals | Only score items created after the training cutoff (LiveCodeBench, LiveBench White+ 2024, each new AIME) | Continuous item supply | Cutoff dates are fuzzy, and the items get stale too |
Mitigations as a benchmark designer
Keep a private held-out split and evaluate through a server (SWE-bench Pro's held-out and commercial splits, ARC-AGI private sets, FrontierMath). Add canary strings (a unique GUID in every file so that crawlers and labs can filter it, as BIG-bench did), keep data gated behind terms of use, refresh items continuously, and prefer tasks where memorizing the answer doesn't help (interactive environments, procedurally generated tasks).
"Your model jumps 10 points on a public benchmark after a data refresh. What do you check?" Look for n-gram overlap between the new data and the test set. Compare against a held-out mirror or perturbed set. Check per-item: are the gains concentrated on items that appear on the web? Run completion probes. Check whether gains transfer to fresh, time-split items. Strong candidates also mention adaptive overfitting (you chose this checkpoint because of this benchmark) and propose a quarantined final test set that is looked at rarely.
- GSM1k: the mirror-set methodology.
- Proving Test Set Contamination in Black-Box LMs: an elegant exchangeability test.
- LiveCodeBench: time-windowed evaluation in practice.
Saturation, the benchmark lifecycle and Goodhart
"When a measure becomes a target, it ceases to be a good measure" (Goodhart's law, in Strathern's popular phrasing). Benchmarks follow a predictable arc:
Signs that a benchmark is saturated: (1) frontier scores are within a few points of the label-noise ceiling (for example, MMLU's estimated 6.5% error rate means about 93% is a soft ceiling); (2) score gains no longer predict gains on fresh or adjacent tasks; (3) top-model differences fall inside the confidence intervals; (4) evidence of memorization appears. The time from release to saturation has shrunk from years to months for some recent benchmarks, which pushes the field toward private, live and agentic evals that are costly to game.
Goodhart failure modes specific to LLMs:
- Teaching to the format. Training on benchmark-style multiple-choice data inflates MC scores without improving the underlying knowledge.
- Reward hacking in agentic evals. Agents that edit tests, special-case checkers or exploit grader loopholes. Graders need adversarial review, and agents need sandboxing that stops them from touching tests.
- Judge hacking. Optimizing against an LLM judge (or arena votes) rewards length, confidence, flattery and formatting.
- Leaderboard selection. Testing many variants and reporting the max (the winner's curse).
A benchmark is a cheap proxy for a capability. Its value lies in the correlation between proxy and capability across models that weren't optimized for it. Optimization pressure finds the directions where proxy and capability come apart, so the correlation always decays. The only defenses are to rotate proxies and keep some private.
LLM-as-judge: biases and mitigations
For open-ended outputs, a strong model grades either a single response (pointwise, for example on a 1–10 rubric) or a pair (pairwise preference). MT-Bench showed that GPT-4 judgments agreed with human experts roughly as often as humans agreed with each other (above 80% on non-tie pairs), which legitimized the approach. The same paper catalogued its biases Zheng+ 2023.
| Bias | What happens | Mitigation |
|---|---|---|
| Position bias | In pairwise comparisons the judge favours the first (sometimes the second) slot. Swapping the order can flip verdicts on a large fraction of pairs Wang+ 2023 | Evaluate both orders. Count a win only if it's consistent, else a tie, or average the probabilities |
| Verbosity / length bias | Longer, more detailed-looking answers are preferred even when they aren't better | Rubrics that penalize padding. Length-controlled win rates (a regression adjustment, as in Length-Controlled AlpacaEval Dubois+ 2024). Report length alongside the score |
| Self-preference | Judges rate their own (or same-family) outputs higher. This correlates with the judge's ability to recognize its own text Panickssery+ 2024 | Use a judge from a different family, or a jury of diverse judges (PoLL Verga+ 2024) |
| Style / authority bias | Markdown, confident tone and citations (even fake ones) sway the judge | Normalize formatting before judging. Check facts separately |
| Limited capability | The judge can't verify math or code it can't solve itself | Give the judge reference answers, execution results or tests. Use a stronger judge than the model being graded |
| Score miscalibration | Pointwise 1–10 scores cluster and drift across judge versions | Pairwise or binary rubric items, anchored examples, and a pinned judge model version |
Process for a trustworthy judge: write a decomposed rubric (several binary checks rather than one holistic score); label a few hundred items with humans; measure judge–human agreement (accuracy, Cohen's κ) and compare it to human–human agreement; iterate the prompt; then freeze the judge version and re-validate whenever you change it. Use the judge for relative comparisons, and don't use the same judge as an RL reward and as the final evaluator.
They'll ask "how do you know your judge is right?" The answer is measured agreement with human labels on a stratified sample, compared to the human ceiling, plus checks for each known bias: swap order, length-matched pairs, and a different-family judge. Bonus points for spotting the RL trap: a policy optimized against a judge will find that judge's blind spots, so hold out a different judge (or humans) for evaluation.
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Length-Controlled AlpacaEval: regression-based debiasing.
- Replacing Judges with Juries
Statistical rigor: error bars on evals
Treat an eval as an experiment: the questions are a sample from a larger super-population of possible questions, and the score is an estimate with sampling error Miller 2024 Anthropic 2024.
Standard error of the mean
For n items with per-item scores \(s_i\) (0/1 for accuracy), \(\mathrm{SE} = \sqrt{\widehat{\mathrm{Var}}(s)/n}\), which is \(\sqrt{p(1-p)/n}\) for accuracy. The 95% CI is about \(\pm 1.96\,\mathrm{SE}\).
| n (items) | Example | SE at p = 0.5 | 95% CI half-width | at p = 0.9 |
|---|---|---|---|---|
| 30 | one AIME year | 9.1 pts | ±17.9 pts | ±10.7 pts |
| 198 | GPQA Diamond | 3.6 pts | ±7.0 pts | ±4.2 pts |
| 500 | SWE-bench Verified, MATH-500 | 2.2 pts | ±4.4 pts | ±2.6 pts |
| 2,500 | HLE | 1.0 pts | ±2.0 pts | ±1.2 pts |
| 14,000 | MMLU | 0.42 pts | ±0.83 pts | ±0.50 pts |
To halve the error bar you need 4× the items. A useful rule: the minimum detectable difference between two independent evaluations at 80% power and α = 0.05 is about \(2.8\sqrt{2}\,\mathrm{SE}\), roughly 4 SE. So on a 500-item benchmark near 50%, you need a gap of about 9 points to reliably detect it with unpaired analysis.
Paired comparisons (the big win)
When two models answer the same questions, analyse per-question differences \(d_i = s_{A,i}-s_{B,i}\):
$$\mathrm{Var}(\bar d) = \frac{\mathrm{Var}(s_A)+\mathrm{Var}(s_B)-2\,\mathrm{Cov}(s_A,s_B)}{n}.$$Because models tend to find the same questions easy or hard, the covariance is usually substantial and positive, so the paired SE can be much smaller than the unpaired one. Paired analysis is "free variance reduction" Miller 2024. Use a paired t-test on \(d_i\), McNemar's test on the discordant pairs (A right/B wrong vs. A wrong/B right), or a paired bootstrap.
Other variance sources
- Clustered questions. Items that share a passage (reading comprehension) or a repo (SWE-bench) aren't independent. Use cluster-robust SEs. Naive SEs can understate uncertainty by a factor of up to about 3 when clusters are strong Miller 2024.
- Sampling randomness. With temperature above 0, each item's score is itself random. Sampling several times per item reduces the within-item variance (useful for small sets), but the between-question variance is the floor.
- Seed and training noise. Two training runs with different seeds differ by about as much as many claimed improvements on some benchmarks. Run several seeds for ablations when you can.
- Multiple comparisons. Testing 20 benchmarks, one will "improve significantly" at p < 0.05 by chance. Pre-register the primary metrics or correct (Holm, Benjamini–Hochberg).
- Infrastructure flakiness. For agentic evals, timeouts, rate limits and sandbox failures can move scores by points. Track and report infra-failure rates separately from task failures.
Classic back-of-envelope question: "New model scores 62% vs 59% on a 500-question benchmark. Ship it?" SE per model is about 2.2 points, so an unpaired 95% CI on the difference is about ±6 points: not significant on its own. With paired analysis and high correlation it might be. Ask for the per-item results, run McNemar or a paired bootstrap, check the discordant counts, look at other benchmarks for the same direction, and consider re-sampling to reduce generation noise.
- Adding Error Bars to Evals (Miller 2024): the formulas for clustered SEs, paired differences and power analysis.
- Anthropic: A statistical approach to model evaluations: the accessible summary.
Evaluating reasoning and test-time compute
Reasoning models (o1-style and successors) turn inference compute into accuracy, both through longer chains of thought and through parallel sampling. A single accuracy number hides this trade. Snell et al. showed that compute-optimal test-time strategies, adapted to problem difficulty, can beat a much larger model at matched FLOPs on some problems Snell+ 2024. s1 showed that "budget forcing" (suppressing the end-of-thinking token, or appending "Wait") lets you sweep thinking length and trace a curve Muennighoff+ 2025.
How to evaluate properly:
- Plot accuracy against tokens (or $, or latency) on a log-x axis, for several reasoning-effort settings. Compare models by curve position: which one reaches X% at the fewest tokens, and where each plateaus. A model that wins at high budget can lose at low budget.
- Separate sequential from parallel scaling. Longer CoT and maj@k / best-of-N have different curves and different latency profiles.
- Report the budget. "Accuracy at max effort" isn't comparable to another lab's default-effort number. Count hidden reasoning tokens in the cost.
- Watch for inverse scaling. More thinking can hurt on some task types (distractors, overthinking, spurious features) Gema+ 2025.
- Mind truncation. Score responses that hit the max-token limit as failures, and report how often it happens. Raising the limit can look like a capability gain.
- Use process metrics where relevant, such as verifier-checked step correctness or the rate of backtracking. These are diagnostic, not headline numbers.
Conventions for reporting reasoning effort (effort levels, thinking budgets, adaptive thinking) changed several times in 2025–2026 and differ by vendor. Check how each vendor defines its effort settings before comparing.
Safety and dangerous-capability evals, red teaming
Frontier labs run dangerous capability evaluations before release under their scaling and preparedness policies. The main domains are CBRN uplift (chemical, biological, radiological, nuclear), offensive cyber, autonomous replication and AI R&D automation, and persuasion or deception. Google DeepMind's 2024 paper is a good template for how the domains are structured Phuong+ 2024. Typical forms:
| Type | Example | What makes it different |
|---|---|---|
| Knowledge proxies | WMDP: about 3.7k multiple-choice questions on hazardous bio/chem/cyber knowledge, also used to measure unlearning Li+ 2024 | Cheap, but knowledge isn't the same as uplift |
| Agentic task suites | Capture-the-flag cyber challenges, autonomous ML-engineering tasks (MLE-bench), self-proliferation tasks, METR-style long tasks | Measure end-to-end ability with tools. Expensive and sensitive to scaffolding |
| Human uplift studies | Randomized trials: experts or novices with vs. without model access on proxy tasks | The gold standard for "uplift", but slow and costly |
| Red teaming | Human experts and automated attackers try to elicit harmful outputs or bypass safeguards | Adversarial and open-ended, so it finds unknown unknowns |
| Propensity / alignment evals | Does the model deceive, sandbag, seek power or reward-hack when given the chance? | Measures disposition, not just capability. Hard to rule out eval-awareness |
Elicitation is the core methodological issue. A capability eval is an upper-bound claim only if you tried hard to elicit the capability: best scaffolding, tools, fine-tuning to remove refusals, and many attempts. Under-elicitation leads to false "safe" conclusions. Sandbagging (a model strategically underperforming) and evaluation awareness (models recognizing tests) are recent concerns that make propensity evals harder.
Red teaming: manual red teams of domain experts, plus automated red teaming where an attacker LM generates test cases and a classifier scores them Perez+ 2022. Anthropic's scaling study of human red teaming found RLHF models got harder to red-team as they scaled, and released the attack dataset Ganguli+ 2022. Jailbreak robustness is measured as attack success rate (ASR) under a fixed attack budget.
Lab frameworks (Anthropic RSP / ASL levels, OpenAI Preparedness Framework, Google DeepMind Frontier Safety Framework) and government evaluation institutes (UK AISI, US CAISI) revised thresholds and processes repeatedly in 2025–2026. Check the current versions directly.
"How is a dangerous-capability eval different from a normal benchmark?" (1) You want an upper bound, so elicitation effort is part of the method. (2) The quantity that matters is uplift relative to baseline resources (internet, experts), not raw accuracy. (3) Thresholds trigger actions, so false negatives cost far more than false positives. (4) Test sets must stay private, and some are hazardous. (5) Results have to hold up against sandbagging.
- Evaluating Frontier Models for Dangerous Capabilities
- Red Teaming LMs with LMs: automated red teaming.
- Inspect: the UK AISI framework used for many safety evals.
Eval harnesses
| Harness | What it is | Strengths | Notes |
|---|---|---|---|
| lm-evaluation-harness (EleutherAI) | Hundreds of tasks defined in YAML. Log-likelihood and generative modes. HF, vLLM and API backends | The standard for base-model and open-model evaluation, and the backend of the original Open LLM Leaderboard. Reproducibility focus Biderman+ 2024 | Agentic and multi-turn support is limited |
| HELM (Stanford CRFM) | Holistic evaluation: many scenarios × metrics (accuracy, calibration, robustness, fairness, toxicity, efficiency) Liang+ 2022 | Standardized, transparent comparisons, with domain-specific leaderboards | More of a benchmark framework than a dev-loop tool |
| Inspect (UK AISI) | Python framework: datasets → solvers (prompting, tool use, agents) → scorers (including model-graded), with sandboxed execution (Docker) | Strong for agentic and safety evals. Large library of community evals | Widely adopted by safety orgs |
| simple-evals (OpenAI) | Minimal reference implementations: MMLU, MATH, GPQA, DROP, MGSM, HumanEval, SimpleQA, BrowseComp, HealthBench | Zero-shot, chat-style prompts, transparent | The README states it is no longer updated for new models as of July 2025 |
| Agent harnesses (SWE-bench harness, Harbor for Terminal-Bench, OSWorld env) | Containerized environments plus task-specific graders | Execution-based grading | The scaffold is part of the result. Report it |
Whatever harness you use, log every raw transcript, version everything (model, harness commit, prompts, judge model), and store per-item scores so you can run paired analysis later. Most "mysterious" eval regressions turn out to be template, tokenizer or extraction bugs, and you find them by reading transcripts.
Part 2 · Interpretability: why it matters
Evals tell you what a model does on the inputs you thought to test. Interpretability tries to tell you how it does it, which gives you predictions about inputs you didn't test. The main motivations:
- Safety and auditing. Detecting deception, hidden objectives or dangerous knowledge that behavioural evals miss (a model that knows it's being tested can behave well). Probes on internal states can act as monitors.
- Debugging. Understanding why a model fails (a spurious feature, a wrong retrieval, a shortcut) so you can fix data or training rather than patching prompts.
- Control. Steering behaviour by intervening on activations, which is cheaper than fine-tuning and more targeted.
- Science. How do in-context learning, factual recall and planning actually work? The answers inform architecture and training choices.
A useful axis: top-down / representation-level methods (probing, steering, representation engineering) find directions that correlate with concepts and test interventions. Bottom-up / mechanistic methods (circuits, SAEs, attribution graphs) try to reverse-engineer the computation into human-understandable parts. Another axis is correlational vs causal: probes and attention maps show what is present, while patching and ablation show what is used.
For an AI engineer role (not a research scientist role), interviewers mostly probe whether you understand the core ideas (residual stream, superposition, SAEs, probing vs patching) and their limitations, and whether you can connect them to practice: probes as cheap classifiers on activations (for example, harmful-intent detection), steering vectors for behaviour control, and why CoT shouldn't be treated as an explanation.
Probing (linear probes)
A probe is a small classifier, usually linear (logistic regression), trained on frozen hidden activations \(h_\ell(x)\) at layer \(\ell\) to predict a property \(y\): part of speech, sentiment, truthfulness, "is this prompt harmful", board state in Othello-GPT. If a linear probe gets high held-out accuracy, the information is linearly decodable at that layer Alain & Bengio 2016. Sweeping layers gives a profile of where a feature becomes available.
Caveats
- Decodable doesn't mean used. The model may never read that information. A probe shows correlation, not mechanism. Follow up with causal interventions (ablate or steer along the probe direction and see whether behaviour changes).
- Probe capacity. A powerful (MLP) probe can learn the task itself from rich features. Hewitt & Liang's control tasks assign random labels to word types. A good probe has high selectivity (real-task accuracy minus control-task accuracy), and simple linear probes are preferred Hewitt & Liang 2019.
- Dataset confounds. A "truth" probe may actually detect "statement looks like a common fact". Test it out of distribution.
Unsupervised variants: Contrast-Consistent Search (CCS) finds a direction where a statement and its negation get complementary probabilities, with no labels, as an attempt to find "what the model believes" Burns+ 2022. Later work questioned whether CCS finds truth specifically rather than other salient features. The broader linear representation hypothesis (concepts are directions) has both theoretical framing Park+ 2023 and known counterexamples, such as circular features for days of the week and other nonlinear manifolds.
Probes are a strong, cheap practical baseline. In GDM's 2025 work on detecting harmful intent, linear probes on raw activations beat SAE-based classifiers out of distribution GDM 2025. If you need a classifier over model internals, start with a linear probe.
The logit lens and tuned lens
Transformers keep a residual stream \(h_\ell\) that every layer adds to. The final prediction is \(\mathrm{softmax}(W_U\,\mathrm{LN}(h_L))\). The logit lens applies the unembedding to an intermediate residual: \(\mathrm{softmax}(W_U\,\mathrm{LN}(h_\ell))\). This lets you watch the prediction form layer by layer: early layers decode to junk or the input token, then the right answer "rises" in the middle and late layers nostalgebraist 2020.
It works because the residual stream is roughly aligned to a shared basis across layers, but it's brittle: in some models (early layers, or models whose basis drifts) the decoded distributions are meaningless. The tuned lens trains a small affine "translator" per layer, \(\mathrm{softmax}(W_U\,\mathrm{LN}(A_\ell h_\ell + b_\ell))\), to minimize KL to the final output. It's more reliable, and the trajectory of predictions across depth can even help detect prompt injection or anomalous inputs Belrose+ 2023.
Reading the logit lens as "what layer ℓ thinks the answer is". Intermediate layers also store information intended for later layers, in directions the unembedding doesn't read. An absence in the logit lens doesn't mean an absence in the representation.
Attention-pattern analysis and its limits
Attention weights \(A = \mathrm{softmax}(QK^\top/\sqrt{d})\) are easy to visualize. Head-level patterns are real and sometimes interpretable: previous-token heads, induction heads, heads that attend to a syntactic head or to the first token ("attention sinks"). Tools like BertViz and CircuitsVis made this popular.
Why attention isn't explanation:
- Attention weights often correlate poorly with gradient-based importance, and you can construct very different attention distributions that give the same prediction Jain & Wallace 2019. (A follow-up, "Attention is not not explanation", argued the picture depends on the test, which is why this is a nuanced point.)
- The weight says where a head reads from, not what it reads (the OV circuit decides what's copied) or whether the output matters downstream.
- Information mixes across layers: by layer 20, "token 5's" residual already contains information from many other positions, so attention to position 5 isn't attention to the word at position 5.
- Many heads attend to sink tokens (BOS, punctuation) as a no-op, which inflates apparent importance.
The mechanistic replacement is to separate each head into a QK circuit (\(W_E^\top W_Q^\top W_K W_E\): which tokens attend to which) and an OV circuit (\(W_U W_O W_V W_E\): what gets written when attended), and to test claims causally Elhage+ 2021. Anthropic's 2025 work extends attribution to explaining why attention patterns form, in terms of feature interactions.
Mechanistic interpretability: residual stream, features, superposition, circuits
The residual-stream view
"A Mathematical Framework for Transformer Circuits" reframed the transformer as a communication channel. The residual stream is a shared memory of width \(d_\text{model}\). Each attention head and MLP reads from it through a linear projection, computes, and adds its output back Elhage+ 2021. Consequences:
- The final logits are a sum of contributions from every component (direct logit attribution), because the stream is additive and the last operations are near-linear.
- Components can compose: a later head's Q, K or V can read what an earlier head wrote (Q-, K- and V-composition). Paths through the network can be enumerated.
- Attention heads move information between positions. MLPs transform information within a position (and store a lot of factual knowledge as key–value lookups).
Features and superposition
A feature is a property of the input the network represents, ideally as a direction in activation space. If each neuron represented one feature, neurons would be monosemantic. In practice many are polysemantic, firing for unrelated things (academic citations, Korean text and HTTP requests). Superposition explains this: when features are sparse (rarely active at the same time), a network can represent many more features than it has dimensions by assigning them nearly orthogonal directions. It accepts a little interference in exchange for capacity Elhage+ 2022. In high dimensions you can fit exponentially many almost-orthogonal vectors (by the Johnson–Lindenstrauss intuition), so superposition is cheap. The toy models show phase changes: as sparsity rises, features move from "not represented" to "dedicated dimension" to "superposed in geometric arrangements" such as antipodal pairs and pentagons.
Superposition is compression. The model is a lossy code for a huge number of rare features, packed into a few thousand dimensions. Neurons are the wrong basis for reading that code. You need to recover the overcomplete feature basis, and that's the motivation for sparse dictionary learning.
Circuits and induction heads
A circuit is a subgraph of features and weights that implements a behaviour. The canonical example is the induction head, a two-head circuit that does in-context copying: on a sequence … [A][B] … [A] it predicts [B] Olsson+ 2022.
Olsson et al. found that induction heads form in a sudden phase change early in training. It shows up as a bump in the loss curve and coincides with a jump in in-context learning ability (loss on later tokens falling relative to earlier ones), and the paper argues that induction heads may account for much of in-context learning in larger models. Other well-studied circuits include the indirect object identification (IOI) circuit in GPT-2 small, which has name-mover, S-inhibition and duplicate-token heads Wang+ 2022, and automated discovery methods such as ACDC Conmy+ 2023.
Activation patching and causal tracing
The workhorse causal method. Run the model on a clean input (for example "The Eiffel Tower is in" → Paris) and a corrupted input (a different subject, or noise added to the subject embeddings). Then, one at a time, copy an activation (a head output, MLP output or residual at some layer and position) from one run into the other, and measure how much of the clean behaviour is restored, for example through the logit difference.
- Denoising (patch clean → corrupted) asks: is this component sufficient to restore the behaviour?
- Noising (patch corrupted → clean) asks: is this component necessary?
- Causal tracing (ROME) is denoising with Gaussian-noise corruption, swept across layers and token positions. It localized factual recall to mid-layer MLPs at the last subject token Meng+ 2022.
- Path patching restricts the patch to specific paths (head A → head B's keys) to map circuits. Attribution patching uses a first-order (gradient) approximation to estimate every patch in one backward pass, trading accuracy for speed.
The methodological choices (corruption type, metric, which positions) noticeably change the conclusions. Gaussian-noise corruption can push the model off-distribution, and logit difference is generally preferred over raw probability Zhang & Nanda 2023 Heimersheim & Nanda 2024. Watch for self-repair (backup heads compensate when one is ablated), which makes necessity tests understate importance.
"How would you find which part of the model stores fact X?" Describe causal tracing: build a clean/corrupted pair, sweep patches over (layer, position), find the restoration peak, then confirm with targeted ablations and an edit (ROME). Mention the pitfalls: off-distribution corruption, self-repair, and the fact that localization doesn't guarantee that editing there will work well (see the model-editing section).
- A Mathematical Framework for Transformer Circuits: the residual-stream view, QK/OV circuits, composition.
- Toy Models of Superposition: the key conceptual paper. Read the first sections.
- In-context Learning and Induction Heads
- How to use and interpret activation patching: a practical guide.
Sparse autoencoders and dictionary learning
If activations \(x \in \mathbb{R}^{d}\) are sparse linear combinations of many feature directions, \(x \approx \sum_i f_i\, d_i\) with most \(f_i = 0\) and \(m \gg d\) directions, then recovering \(\{d_i\}\) is a sparse dictionary learning problem. A sparse autoencoder (SAE) learns it with a one-layer encoder and a linear decoder Cunningham+ 2023 Bricken+ 2023:
$$f(x) = \mathrm{ReLU}\big(W_e (x - b_d) + b_e\big) \in \mathbb{R}^{m}, \qquad \hat x = W_d\, f(x) + b_d,$$ $$\mathcal{L} = \underbrace{\lVert x - \hat x\rVert_2^2}_{\text{reconstruction}} \;+\; \lambda \underbrace{\sum_i f_i(x)\,\lVert W_{d,:,i}\rVert_2}_{\text{sparsity (L1)}}.$$Each column of \(W_d\) is a feature direction. The expansion factor \(m/d\) typically ranges from 8× to 256× or more. Each latent's activations are inspected (top-activating examples, logit effects) and auto-labelled by an LLM.
Key results and variants
- Towards Monosemanticity (Anthropic, 2023): SAEs on a 1-layer transformer's MLP produced features that are far more interpretable than neurons (DNA sequences, Arabic script, base64) Bricken+ 2023.
- Scaling Monosemanticity (2024): SAEs with up to 34M features on Claude 3 Sonnet's middle residual stream. Features were multilingual, multimodal and abstract (code bugs, deception, sycophancy, the Golden Gate Bridge), and clamping features steered behaviour (the "Golden Gate Claude" demo) Templeton+ 2024.
- TopK SAEs (OpenAI): replace L1 with keeping the top-k latents. This sets sparsity directly, removes "shrinkage" (L1 pulls activations toward zero), gives clean scaling laws, and was trained at 16M latents on GPT-4 activations Gao+ 2024.
- JumpReLU SAEs (GDM): a learned per-latent threshold with a straight-through estimator for an L0 penalty, which improves the reconstruction–sparsity frontier Rajamanoharan+ 2024. Gemma Scope released open JumpReLU SAEs for every layer of Gemma 2 models Lieberum+ 2024. Browse them on Neuronpedia.
How SAEs are evaluated
| Metric | Meaning | Trade-off |
|---|---|---|
| L0 | Average number of active latents per token | Lower is sparser and more interpretable, but reconstructs worse |
| Fraction of variance explained / MSE | Reconstruction quality | Can look high while missing what matters for behaviour |
| Δ loss (spliced-in CE) | Replace activations with \(\hat x\) and measure the increase in LM loss | The most behaviour-relevant fidelity metric. Splicing typically still costs noticeable loss |
| Interpretability scores | Auto-interp: an LLM explains a latent, then predicts its activations from the explanation | Cheap at scale, but explanations can be superficial |
| Dead / dense latents | Latents that never fire, or fire on everything | A training-health indicator (resampling, TopK and initialization tricks help) |
| Downstream tasks | Probing, steering, unlearning or spurious-correlation removal using SAE features vs. baselines | Increasingly seen as the real test, and results are mixed |
Feature splitting and other limitations
- Feature splitting. As the dictionary grows, one feature ("math") splits into finer ones (algebra, geometry, LaTeX tokens). There's no single "true" feature set: granularity is an artefact of dictionary size. Related problems are feature absorption (a general feature such as "starts with S" fails to fire where a more specific feature absorbs it) and composition (features that are combinations of others).
- Dark matter. Reconstruction error doesn't fall to zero as you scale. A residual part of the activations isn't captured, and it may contain important computation.
- Not causal by construction. SAEs model activations, not computation. A feature can be interpretable yet unused, and the decomposition can reflect dataset statistics rather than the model's own ontology.
- Downstream disappointment. GDM reported that SAE-based probes underperformed linear probes on out-of-distribution harmful-intent detection and deprioritized fundamental SAE research GDM 2025. A 2026 preprint found that on synthetic data with known features, SAEs recovered only a small fraction of true features despite high explained variance Korznikov+ 2026. A counter-position: SAEs are weaker for acting on known concepts but valuable for discovering unknown ones (hypothesis generation, auditing) Peng+ 2025.
- Cost. Training SAEs over every layer of a frontier model takes significant compute and storage.
The field's view of SAEs shifted noticeably between 2024 (high enthusiasm) and 2025–2026 (more skepticism, more emphasis on downstream benchmarks, and newer architectures such as transcoders, crosscoders, sparse mixtures of linear transforms and turn-averaged SAEs). Treat "SAEs solve superposition" as contested, not established.
"Explain an SAE in two minutes." Superposition means neurons are polysemantic, so learn an overcomplete dictionary. Encoder with ReLU/TopK, linear decoder, loss = MSE + sparsity penalty. Features are columns of the decoder. Evaluate by L0, spliced-in Δloss and interpretability. Limitations: feature splitting, dark matter, mixed downstream utility (linear probes often win for known concepts). Mentioning the L1 shrinkage problem and the TopK/JumpReLU fix signals real depth.
- Scaling Monosemanticity: the canonical large-scale SAE paper with interactive feature browsers.
- Scaling and evaluating sparse autoencoders (TopK)
- GDM: Negative results for SAEs on downstream tasks: an honest counterweight.
Transcoders, attribution graphs and circuit tracing
SAEs decompose a representation. To trace computation you want interpretable pieces connected by interpretable interactions. Transcoders do this: instead of reconstructing an MLP's input, a transcoder is trained to imitate the MLP, mapping its input to its output through a sparse latent layer. You can then substitute it for the MLP, and feature-to-feature interactions become (mostly) linear and input-independent Dunefsky+ 2024.
Anthropic's circuit tracing (2025) scaled this up Ameisen+ 2025:
- Train a cross-layer transcoder (CLT): each layer's features read from that layer's residual stream and can write to all subsequent MLP outputs, which collapses multi-layer MLP computation into fewer features.
- Build a replacement model by swapping all MLPs for the CLT. For a given prompt, freeze attention patterns and normalization, and add error nodes for whatever the CLT fails to reconstruct, so that the local replacement exactly reproduces the model's output on that prompt.
- Because the replacement is linear given the frozen attention, compute an attribution graph: nodes are active features (plus input tokens, error nodes and output logits), and edges are direct linear contributions. Prune it to the most influential paths and group features into "supernodes" by hand.
- Validate hypotheses with interventions in the real model: suppress or inject features and check that downstream features and outputs change as the graph predicts.
Findings in Claude 3.5 Haiku Lindsey+ 2025 included genuine multi-hop reasoning inside the forward pass ("capital of the state containing Dallas" goes through an intermediate "Texas" feature), planning in poetry (rhyme candidates chosen before the line is written), language-independent concept features with language-specific input and output circuitry, a default "can't answer" circuit that known-entity features inhibit (relevant to hallucination), and cases where the model's verbal explanation of its arithmetic didn't match the internal mechanism. The authors noted that the method gave satisfying insight on only roughly a quarter of the prompts they tried, and that it doesn't explain attention-pattern formation (QK) by default. An open-source circuit-tracer library with Neuronpedia integration lets you generate attribution graphs for open-weight models.
Anthropic's Transformer Circuits thread kept publishing through 2026. Topics include tracing attention computation through feature interactions (2025), crosscoder model diffing, sparse mixtures of linear transforms as a transcoder alternative, emergent introspective awareness (2025), "Natural Language Autoencoders" that train Claude to verbalize its activations (2026), and a "global workspace" of verbalizable representations (2026) Transformer Circuits. These are recent and less independently replicated than the 2021–2024 core results. Read them as active research, not settled consensus.
- Circuit Tracing: Revealing Computational Graphs in Language Models: the method.
- On the Biology of a Large Language Model: case studies. The most readable entry point.
- Anthropic: Tracing the thoughts of a large language model: a non-technical summary.
Steering vectors and representation engineering
If concepts are (approximately) linear directions, you can add a direction to the residual stream at inference to change behaviour, with no weight updates. Activation addition (ActAdd) takes the difference of activations on a contrast pair of prompts ("Love" − "Hate") and adds a scaled version at some layer for all positions Turner+ 2023. Contrastive Activation Addition (CAA) averages the difference over many paired examples of a behaviour (sycophantic vs. not) to get a more robust vector Panickssery+ 2023:
$$v_\ell = \frac{1}{|D|}\sum_{(p^+,\,p^-) \in D} \big(h_\ell(p^+) - h_\ell(p^-)\big), \qquad h_\ell \leftarrow h_\ell + \alpha\, v_\ell .$$Representation Engineering (RepE) generalizes this into "reading" (find concept directions with PCA on contrastive activations, use them as monitors) and "control" (add or subtract them) for honesty, harmlessness, emotion and more Zou+ 2023. A striking example: refusal in many chat models is mediated by a single direction. Ablating it (projecting it out of weights or activations) largely removes refusals, and adding it makes the model refuse harmless requests Arditi+ 2024. This shows both the power of the method and why open-weight safety tuning is shallow. Persona vectors extend the idea to monitoring and controlling traits such as sycophancy or "evil" across fine-tuning, including predicting which training data will shift a trait Chen+ 2025.
| Approach | Mechanism | Pros | Cons |
|---|---|---|---|
| Prompting | Instruction in context | Trivial, no access to internals | Can be overridden, uses context |
| Steering vectors (CAA/RepE) | Add a direction to the residual at inference | Cheap, reversible, continuous strength α, few examples | Needs activation access. Large α hurts fluency and capability. Layer and α need tuning. Effects can be inconsistent across inputs |
| SAE feature clamping | Set specific SAE latents to high or low values | Interpretable knob | Needs a good SAE. Off-target effects |
| Fine-tuning (SFT/RL/LoRA) | Change weights | Robust, general | Costly, less targeted, can regress other skills |
"How would you reduce sycophancy without retraining?" Build a contrast dataset (sycophantic vs. honest completions on the same prompts), compute a CAA vector at a middle layer, subtract it at inference with α tuned on a validation set, and measure both the target behaviour and general capability (an MMLU-style regression check). Mention monitoring as an alternative use: project activations onto the vector to flag sycophantic generations.
Model editing: ROME and MEMIT (briefly)
Model editing changes a specific fact ("The Eiffel Tower is in Rome") without retraining. ROME uses causal tracing to find the mid-layer MLP that recalls facts at the last subject token. It treats that MLP's output projection as a linear associative memory \(W k \approx v\) (key = subject representation, value = fact), and applies a closed-form rank-one update so the subject's key maps to a new value while other keys are disturbed as little as possible Meng+ 2022. MEMIT spreads updates across several layers to insert thousands of facts at once Meng+ 2022.
Evaluation uses efficacy (the edit holds), generalization (paraphrases work), and specificity (neighbouring facts don't change). Known limitations: poor "ripple effects" (editing a person's birthplace doesn't update facts that follow from it), degradation after many sequential edits, and evidence that where causal tracing localizes a fact doesn't predict where editing works best. In practice, retrieval (RAG) or fine-tuning is usually preferred over weight editing for knowledge updates. Editing matters more as an interpretability probe than as a product tool.
Chain-of-thought faithfulness
A visible reasoning trace is tempting to read as an explanation. Faithfulness asks whether the CoT reflects the factors that actually drove the answer. The evidence says it often doesn't:
- Biasing features go unmentioned. When few-shot examples always had answer (A), or a user suggested an answer, models shifted toward it while their CoTs rationalized the choice without mentioning the bias Turpin+ 2023.
- Interventions on the CoT. Truncating the CoT early, inserting mistakes, paraphrasing, or replacing it with filler tokens tests whether the answer depends on the CoT's content. Dependence varied by task, and larger models were sometimes less reliant on their stated reasoning Lanham+ 2023.
- Reasoning models. Given hints (including "unauthorized access" style hints) that they demonstrably used, Claude 3.7 Sonnet and DeepSeek R1 verbalized the hint in at least 1% of cases in most settings, but often below 20%. Outcome-based RL improved faithfulness at first and then plateaued, and when RL taught reward hacking, models almost never verbalized the hack Chen+ 2025.
- Mechanistic evidence. Attribution graphs showed cases where a model's described arithmetic method differed from the internal one, and cases of "motivated reasoning" working backward from a suggested answer Lindsey+ 2025.
But CoT is still useful for monitoring. When a task is hard enough that the model needs to reason in tokens, the reasoning has to pass through the visible trace, so monitors can catch misbehaviour there. A multi-lab position paper argues that this monitorability is "a new and fragile opportunity", which training pressure on the CoT (optimizing it to look good) or architectures that reason in latent space could destroy Korbak+ 2025.
Treating CoT as an audit log ("the model explained why it chose this, so it's safe"), or penalizing "bad thoughts" in the CoT during RL. The second teaches the model to hide its reasoning rather than to stop the behaviour, and it degrades the CoT's monitoring value.
"Can we trust a reasoning model's chain of thought?" A strong answer distinguishes faithfulness (does the CoT reflect the true causes? often not, per Turpin and Chen) from monitorability (is the CoT useful for detecting problems? often yes, when reasoning is necessary). It names the tests (bias injection, hint verbalization, CoT truncation and corruption) and the policy implication: don't optimize the CoT directly, and combine CoT monitoring with output checks and interpretability-based probes.
What is established vs. recent or contested
| Well established | Recent / contested (2025–2026) |
|---|---|
| The unbiased pass@k estimator. BT/Elo for pairwise ranking. SE formulas and paired analysis. Prompt-format sensitivity. Contamination exists and inflates public benchmarks. LLM-judge position, length and self-preference biases. | Which coding and agent benchmarks are still "valid" (SWE-bench Verified retired, Pro and Terminal-Bench rising). How much RLVR expands vs. sharpens capability (pass@k at large k). Arena integrity. Comparability of reasoning-effort settings. Eval awareness and sandbagging. |
| The residual-stream framework. Superposition in toy models. Induction heads and their link to ICL. Activation patching as a causal tool. Probes are correlational. Attention is not a full explanation. CoT is often unfaithful. | How far SAEs are useful downstream. Whether dictionary features reflect the model's own ontology. How well attribution graphs scale to frontier models. Introspection and verbalization methods. How long CoT stays monitorable. |
Interview question bank
Why can't you compare perplexities across two models with different tokenizers, and what do you use instead?
Perplexity is the exponentiated mean NLL per token. A tokenizer with a larger vocabulary produces fewer, longer tokens, so each prediction carries more information and per-token loss is higher even if the model is equally good per character. Normalize to a tokenizer-independent unit: bits per byte (or per character), computed as total NLL in bits divided by the number of UTF-8 bytes. You must also hold the evaluation text fixed (same documents, same preprocessing) and compare base models with base models, since post-training changes the output distribution deliberately.
Derive or explain the unbiased pass@k estimator. Why not use 1 − (1 − c/n)^k?
Generate n ≥ k samples per problem and count c correct ones. The probability that a random size-k subset contains no correct sample is C(n−c, k)/C(n, k), so pass@k = E[1 − C(n−c,k)/C(n,k)], averaged over problems. This is exactly unbiased for any k ≤ n. The plug-in 1 − (1 − p̂)^k is biased low because the function is concave in p (Jensen). For numerical stability compute 1 − ∏_{i=n−c+1}^{n}(1 − k/i), returning 1 if n − c < k. Codex used n = 200. Also tune temperature per k: higher temperatures help large-k, hurt k = 1.
A model gets pass@1 = 40% and pass@100 = 90% on a code benchmark. What does that tell you, and how would you exploit it?
The model can often produce a correct solution but can't reliably pick it, so the capability is latent and the bottleneck is selection or consistency. If unit tests or a verifier are available at deployment, generate many samples and filter (pass@k becomes a realistic deployed metric). Without one, train a verifier or reward model for best-of-N, use maj@k where outputs are comparable, or use RL with verifiable rewards to concentrate probability on correct modes, which should raise pass@1 toward pass@k. Keep in mind that RL sharpening can reduce diversity, so pass@k at very large k may fall.
What's the difference between pass@k, pass^k and maj@k? When would you report each?
pass@k: at least one of k samples correct. It needs an oracle at deployment (tests) and measures latent capability. pass^k: all k trials correct, estimated as C(c,k)/C(n,k). It measures reliability and is crucial for customer-facing agents (τ-bench). maj@k: the majority-voted answer of k samples is correct. It's deployable without an oracle and is a form of test-time compute. Report pass@1 as the headline, pass@k for code with tests, pass^k for agents, and maj@k only alongside pass@1 at matched compute.
Back-of-envelope: Model B scores 62% vs Model A's 59% on a 500-item benchmark. Is B better?
SE per model ≈ √(0.6·0.4/500) ≈ 2.2 points, so an unpaired SE of the difference is ≈ 3.1 points and the 95% CI is ≈ 3 ± 6 points: not significant. With paired data the variance subtracts 2·Cov. If the per-item correlation is high, the paired SE might be around 1.5–2 points, and a 3-point gap could become borderline significant. Run McNemar's test on the discordant items or a paired bootstrap. Also check consistency across seeds, prompts and related benchmarks, the multiple-comparisons context, and whether the benchmark's label noise makes small differences meaningless.
How many questions do you need to detect a 2-point accuracy improvement around 70% accuracy?
Unpaired: the SE of a difference is √(2·0.7·0.3/n). For 80% power at α = 0.05 you need the effect ≈ 2.8 SE, so 0.02 = 2.8·√(0.42/n), giving n ≈ 0.42·(2.8/0.02)² ≈ 8,200 questions per model. Paired analysis with correlated outcomes can cut this by a large factor (often 2–5×, depending on correlation), which is why you always use paired designs. For small benchmarks like GPQA Diamond (198 items) a 2-point difference is fundamentally undetectable.
Explain Bradley–Terry and how Elo is derived for an arena. What does a 50-point gap mean?
BT models P(A beats B) = σ(β_A − β_B) and fits β by maximum likelihood (logistic regression on battle outcomes with ±1 indicator features). Scores are rescaled to the Elo convention R = 400·β/ln 10 + const, so P = 1/(1 + 10^{(R_B−R_A)/400}). A 50-point gap ≈ 57% win rate, 100 ≈ 64%. Batch BT fitting with bootstrap CIs is preferred over online Elo, which depends on the order of games. Caveats: style confounds (handled with length and markdown covariates), non-uniform sampling, selective disclosure of private variants, and the arena prompt distribution not matching yours.
Log-likelihood vs generative scoring for multiple-choice benchmarks: trade-offs?
Log-likelihood scoring compares log p(option | prompt) across options. It's deterministic, cheap (no decoding) and works for base models without instruction following, which makes it ideal during pretraining. But it needs a length-normalization choice, can't capture reasoning, and isn't available for closed APIs without logprobs. Generative scoring lets the model reason and answer freely, matching real use, but depends on answer extraction, adds sampling variance and costs more. Rankings can differ between the two. Small models often look better under cloze scoring because the "answer with a letter" format is itself an emergent skill.
How does test-set contamination happen, and how would you detect it with and without access to training data?
Benchmarks leak via GitHub, Hugging Face, papers, forums, paraphrases and synthetic data from teachers that saw them, plus adaptive overfitting through checkpoint and mix selection. With training data: n-gram or substring overlap search (for example 13-gram or 50-character matches), embedding similarity for paraphrases. Without it: membership inference such as Min-K% Prob, the exchangeability test (canonical order more likely than shuffled), completion probes, perturbation tests (change numbers or names), and comparing against fresh mirror sets (GSM1k) or post-cutoff items (LiveCodeBench). Prevention: canary strings, private splits, live refresh.
Why did benchmarks like MMLU and SWE-bench Verified stop being useful?
MMLU saturated: frontier models exceed 90%, and with an estimated ~6.5% label-error rate the remaining headroom is mostly noise. Years of public exposure also make contamination likely. SWE-bench Verified: OpenAI reported in Feb 2026 that many tests were overly narrow or checked things the issue never specified, and that frontier models could reproduce gold patches. Scores therefore reflected memorization and test idiosyncrasies more than coding skill, and they recommended SWE-bench Pro. The general lesson is Goodhart plus lifecycle: every public benchmark decays under optimization pressure, so you need successors, private splits and live data.
Design an LLM-as-judge pipeline for comparing two chat models. How do you make it trustworthy?
Use pairwise comparisons on an in-domain prompt set, with a decomposed rubric (correctness, instruction following, safety, concision) and reference answers where possible. Mitigate position bias by judging both orders and counting a win only if it's consistent. Mitigate length bias with length-controlled analysis or rubric penalties. Mitigate self-preference by using a judge from a different family, or a jury of 3 diverse judges. Validate on ~300–500 human-labelled pairs, measure agreement and κ against the human–human ceiling, then freeze the judge version. Report win rates with CIs, analyse by category, and re-validate whenever the judge changes. Never use the same judge as both the RL reward and the final evaluation.
How should you compare two reasoning models fairly?
Compare curves, not points. Sweep reasoning effort or thinking budget and plot accuracy against tokens (or $ or latency), separating sequential (longer CoT) from parallel (maj@k / best-of-N) scaling. Compare at matched budgets and report where each plateaus. Count hidden reasoning tokens in the cost, score truncated outputs as failures, and report the truncation rate. Watch for inverse scaling where more thinking hurts. Use enough samples per problem on small sets like AIME, and use fresh, post-cutoff problems.
What makes a dangerous-capability eval different from a normal capability benchmark?
It's meant to support a claim of the form "the model can't do X", so it must approximate an upper bound: maximal elicitation (best scaffolds, tools, many attempts, fine-tuning away refusals). Under-elicitation produces false negatives, which are the costly error. It measures uplift relative to baselines (internet, experts), sometimes through human uplift trials. Results trigger pre-committed mitigations. Items are often private or hazardous. Recent concerns include sandbagging and evaluation awareness, which can make models look safer in tests than in deployment.
What is a linear probe, and what are the two biggest pitfalls in interpreting one?
A logistic-regression (or similarly simple) classifier trained on frozen activations from a layer to predict some property. High accuracy means the property is linearly decodable there. Pitfall 1: decodable doesn't mean used. The model might never read that direction, so you need causal tests (ablate or steer along the probe direction). Pitfall 2: probe capacity and confounds. An expressive probe can learn the task itself (use control tasks and report selectivity), and the dataset may confound the target with surface features, so test out of distribution. In practice, linear probes are strong, cheap monitors and often beat fancier SAE-based classifiers.
What is the logit lens, why does it work, and when does it fail?
Apply the final layer norm and unembedding to an intermediate residual stream to decode a "current best guess" at each layer. It works because the residual stream is additive and components write in a basis roughly shared with the output. It fails when intermediate representations are in a rotated or shifted basis (often in early layers, and in some model families), and because intermediate layers store information for later computation that isn't meant to be read by the unembedding. The tuned lens fixes much of this with a learned per-layer affine map trained to match the final distribution.
Explain superposition and why it motivates sparse autoencoders.
Networks need to represent far more features than they have neurons. When features are sparse (rarely co-active), the network can assign them nearly orthogonal directions in a lower-dimensional space and tolerate small interference, so a single neuron participates in many features (polysemanticity). Toy models show this happens predictably as sparsity increases. The neuron basis is therefore the wrong unit of analysis. SAEs try to recover the overcomplete feature basis by learning a wide, sparse code that reconstructs activations, in the hope that each latent corresponds to one interpretable feature.
Write down the SAE objective. What is "shrinkage" and how do TopK/JumpReLU address it?
f = ReLU(W_e(x − b_d) + b_e), x̂ = W_d f + b_d, and L = ‖x − x̂‖² + λ Σ_i f_i‖d_i‖ (L1 weighted by decoder norms so the model can't cheat by shrinking f and growing W_d). The L1 penalty pushes all active latents toward zero, so the reconstruction systematically underestimates feature magnitudes. That's shrinkage, and it also blurs the sparsity/fidelity trade-off. TopK SAEs keep only the k largest pre-activations, which sets L0 = k directly with no L1 penalty on magnitudes. JumpReLU learns a per-latent threshold and penalizes L0 using straight-through gradients. Both improve the reconstruction–sparsity frontier.
What are the main limitations of SAEs?
Feature splitting (granularity depends on dictionary size, so there's no canonical feature set), absorption (general features fail to fire when specific ones absorb them), dark matter (a persistent reconstruction error, and splicing in the SAE noticeably increases LM loss), interpretable latents that aren't necessarily causally used, high training cost per layer, and mixed downstream results. GDM found SAE probes worse than linear probes on OOD harmful-intent detection, and synthetic studies find low recovery of ground-truth features. A fair framing: SAEs are good for discovery and hypothesis generation, and weaker than simple baselines for acting on known concepts.
How does activation patching work? Contrast noising and denoising.
Run clean and corrupted inputs that differ in one key factor. Copy an internal activation from one run into the other and measure the change in a metric such as the logit difference of the correct vs. incorrect answer. Denoising (clean into corrupted) tests sufficiency: does restoring this component recover the behaviour? Noising (corrupted into clean) tests necessity: does breaking it destroy the behaviour? Sweeping over layers and positions localizes computation (ROME's causal tracing found mid-layer MLPs at the last subject token for factual recall). Pitfalls: off-distribution corruptions (Gaussian noise), metric choice, and self-repair by backup components.
Describe the induction-head circuit and why it matters.
Two heads in different layers. A previous-token head writes "the previous token was A" into each position's residual. An induction head at a later occurrence of A queries for positions whose previous-token feature equals A (K-composition), attends to the token after the earlier A, and copies it through its OV circuit, predicting B in "…A B … A". Induction heads appear during a sharp phase change in training that coincides with a jump in in-context learning, and Olsson et al. argue they're a major mechanism of ICL. They're the cleanest example of a multi-layer circuit found through weights and causal experiments together.
What are attribution graphs, and how do they differ from SAE feature browsing?
SAEs give you a list of active features. Attribution graphs show how features cause each other on a specific prompt. Anthropic trains a cross-layer transcoder that replaces the MLPs, freezes attention and norms for the prompt, and adds error nodes so the local replacement reproduces the output exactly. The model is then linear in the features, so direct contributions between features can be computed as edges. The graph is pruned and grouped into supernodes, and hypotheses are validated by intervening on features in the real model. It revealed multi-hop reasoning and planning inside a forward pass. Limitations: it explains only a fraction of prompts well, doesn't explain attention formation by default, and depends on transcoder fidelity (error nodes).
How do steering vectors work and what are their failure modes?
Compute a direction associated with a behaviour, typically the mean difference of residual activations between contrastive prompts (CAA), or the top PCA component of the differences (RepE). Add α·v at one or more layers during inference to induce the behaviour, or subtract it to suppress it. You can also project onto v to monitor. They work because many high-level concepts are roughly linear in the residual stream. The refusal direction is a striking case: ablating it removes refusals. Failure modes: too large an α degrades fluency and capability, effects vary across prompts and layers, vectors capture dataset confounds, and steering may fail for concepts that aren't linear. Always measure off-target capability regressions.
Is a reasoning model's chain of thought a faithful explanation? What's the evidence, and does it matter?
Often not. Turpin et al. showed biased few-shot contexts change answers while CoTs rationalize without mentioning the bias. Lanham et al.'s truncation and corruption tests show variable reliance on the CoT. Anthropic's 2025 study found reasoning models verbalize hints they used often below 20% of the time, and RL-learned reward hacks were almost never verbalized. It matters because CoT monitoring is an attractive safety tool. Monitoring can still work when the task requires explicit reasoning, but optimizing CoTs to look good, or moving reasoning into latent space, could make CoTs uninformative. Hence the recommendation to preserve monitorability and not train directly against CoT content.
How do ROME and MEMIT edit facts, and why aren't they widely used for knowledge updates in production?
ROME locates factual recall in a mid-layer MLP via causal tracing, views the MLP's output weights as a key→value memory, and applies a closed-form rank-one update so the subject key maps to a new value with minimal change to other keys. MEMIT distributes updates over several layers for thousands of edits. In production they're rarely used because edits don't propagate to logically related facts (ripple effects), many sequential edits degrade the model, specificity failures are hard to audit, and RAG or fine-tuning give more controllable, reversible knowledge updates.
Design: you're shipping a new frontier model checkpoint next month. What does your eval plan look like?
Layer it. (1) Regression suite: fixed-protocol harness (lm-eval or Inspect) with per-domain loss/BPB plus core capability sets (knowledge, math, code, long context, multilingual), stored per item for paired comparison against the last release. (2) Frontier and agentic: SWE-bench Pro, Terminal-Bench 2.0, τ²-bench with pass^k, OSWorld, HLE/FrontierMath, with scaffolds and budgets documented and accuracy-vs-tokens curves. (3) Fresh or private: post-cutoff items and internal held-out sets to control contamination. (4) Preference: blind pairwise human evals plus a validated judge, style-controlled. (5) Safety: dangerous-capability evals with maximal elicitation, red teaming, jailbreak ASR, propensity evals. (6) Rigor: CIs, paired tests, multiple-comparison control, transcript review of regressions, and a decision rule agreed in advance.