Track A · Model internals

Model Evaluation & Interpretability

Two ways to answer "what does this model actually do?" Evaluation looks from the outside: what you measure, how you score it, and how much to trust the number. Interpretability looks from the inside: which internal representations and circuits produce the behaviour. This page treats evaluation the way a model builder sees it (pretraining and post-training checkpoints, public benchmarks, leaderboards, capability and safety evals). Product evals for LLM applications (task rubrics, golden sets, online metrics) are covered separately in B5. Interviewers use this area to tell apart people who read leaderboards from people who can say why a 2-point gain on a 500-question benchmark may be noise, why a judge model prefers longer answers, or what a sparse autoencoder is optimizing.

TL;DR: the 8–12 things to be able to say out loud

  • Perplexity is \(\exp\) of the mean per-token negative log-likelihood. It is great for comparing checkpoints that share a tokenizer and data distribution, but it doesn't compare across tokenizers (use bits-per-byte) and it predicts downstream task quality only loosely, especially after RLHF.
  • Benchmarks follow a lifecycle: introduced hard, then optimized against, contaminated and saturated within roughly 1–3 years. MMLU, GSM8K and HumanEval are saturated. SWE-bench Verified was retired by OpenAI in early 2026 over flawed tests and contamination. Frontier signal now comes from HLE, FrontierMath, ARC-AGI-2/3, Terminal-Bench 2.0, SWE-bench Pro and agentic or long-horizon evals.
  • pass@k should be computed with the unbiased estimator \(1-\binom{n-c}{k}/\binom{n}{k}\) from \(n \ge k\) samples, not by drawing \(k\) samples once. maj@k (self-consistency) is a different thing: it's a single answer chosen by voting.
  • Arena leaderboards fit a Bradley–Terry model to pairwise votes: \(P(A \succ B)=\sigma(\beta_A-\beta_B)\). A 100-point Elo gap means about a 64% win rate. Arenas are biased by style, by who gets sampled, and by selective disclosure.
  • The scoring protocol changes the number: zero- vs few-shot, prompt format (up to tens of points on older models), log-likelihood vs generative scoring, and answer extraction. Always report the protocol.
  • Contamination: test items leak into pretraining data. You detect it with n-gram overlap, membership inference (Min-K% Prob), exchangeability tests, or by building fresh mirror sets (GSM1k) and live benchmarks (LiveCodeBench, LiveBench).
  • LLM-as-judge has position, verbosity and self-preference biases. Mitigations: swap order and average, control for length, use a judge from a different family or a jury of judges, use reference answers, and calibrate against human labels.
  • Statistics: the standard error of accuracy is \(\sqrt{p(1-p)/n}\). With n=500 and p=0.5 that is ±4.4 points at 95% confidence. With n=30 (one year of AIME) it is about ±18 points. Use paired tests and clustered SEs, and resample when the model is stochastic.
  • Reasoning models should be compared on accuracy-vs-tokens (or cost) curves at matched budgets, not on single numbers.
  • Interpretability toolkit: linear probes (correlational), logit lens and tuned lens, attention patterns (suggestive, not explanations), and activation patching (causal). Mechanistic interp views the model as a residual stream that components read from and write to.
  • Superposition: models represent more features than they have dimensions, which is why neurons are polysemantic. Sparse autoencoders learn an overcomplete sparse dictionary, with loss = reconstruction + λ·sparsity. Transcoders and cross-layer transcoders let you build attribution graphs of computation.
  • Steering vectors add a direction (for example a mean activation difference) to the residual stream to change behaviour. ROME/MEMIT edit facts through rank-one updates to MLP weights. CoT is often not faithful: reasoning models verbalize the hints they actually used less than 20% of the time in many settings.

Part 1 · Model evaluation: why it's harder than it looks

Evaluating a classifier on ImageNet was a solved social problem: one fixed test set, one metric, one protocol. LLM evaluation breaks each of those assumptions:

A model builder runs evals at several points, each with its own goal:

StageWhat you measureTypical toolsMain failure mode
Pretraining (during the run)Held-out loss/perplexity, cheap log-likelihood multiple-choice (HellaSwag, ARC-Easy, MMLU cloze), small generative math/code setsFast in-loop evals every N steps, scaling-law fitsMetrics that are flat until a threshold ("emergence"), which leaves no signal for small-scale ablations
Mid/post-trainingInstruction following, chat quality, math/code (generative), tool use, refusal behaviourGenerative harnesses, LLM judges, internal preference evalsOverfitting to the judge or to public benchmarks, prompt-format sensitivity
Release decisionFrontier capabilities, agentic tasks, human preference, dangerous capabilitiesAgent harnesses (sandboxes), arenas, red teams, third-party evaluatorsElicitation gaps (under-estimating what the model can do), cost
Post-releaseReal usage, regressions, arena rankingsArena-style platforms, telemetry (that's B5 territory)Selection bias, style effects
Interview angle

"How would you evaluate a new base model checkpoint vs. the previous one?" A strong answer separates cheap, low-variance, in-loop signals (held-out loss on several domains, bits-per-byte, log-likelihood multiple-choice) from expensive generative evals. It holds the protocol fixed across checkpoints, reports confidence intervals and paired differences, checks contamination of any new data, and picks metrics that are continuous at the current scale (log-prob of the correct answer rather than exact match) so small ablations still show signal.

Perplexity and its limits

For a held-out token sequence \(x_1,\dots,x_N\), the average negative log-likelihood (cross-entropy, in nats) is

$$\mathcal{L} = -\frac{1}{N}\sum_{i=1}^{N}\log p_\theta(x_i \mid x_{\lt i}), \qquad \mathrm{PPL} = e^{\mathcal{L}}.$$

Intuitively, perplexity is the effective branching factor: PPL = 10 means the model is, on average, as uncertain as if choosing uniformly among 10 tokens. It's the pretraining objective itself, so it's smooth, cheap, low-variance, and it is what scaling laws are fit to (see the scaling-laws page).

Why you can't stop at perplexity

Intuition

Loss is the right training signal because it's dense: every token gives gradient. It's an imperfect evaluation signal because the tasks you care about depend on a tiny number of high-stakes tokens (the final answer, the right API call). Benchmarks exist to put weight on those tokens.

Common mistake

Comparing Llama-family and GPT-family perplexities directly, or comparing PPL between a base and a chat model and concluding that the chat model "got worse". Normalize by bytes and only compare like with like.

Go deeper

The benchmark landscape (as of late 2026)

The table groups widely cited benchmarks by what they claim to measure. "Status" is a rough judgement of whether the benchmark still separates frontier models. Frontier scores move monthly, so treat the numbers as indicative.

May be out of date

Benchmark status changes fast. The saturation notes reflect public reporting up to about September 2026. Leaderboard trackers often disagree by many points because they use different protocols (tools vs no tools, text-only subsets, reasoning effort). Check the official leaderboard before quoting a number.

CategoryBenchmarkWhat it isScoringStatus
KnowledgeMMLU Hendrycks+ 2020About 14k four-way multiple-choice questions over 57 subjectsAccuracy (log-likelihood or generative)Saturated: frontier models above 90%. About 6.5% of questions have errors, per MMLU-Redux Gema+ 2024
MMLU-Pro Wang+ 2024Harder, reasoning-heavier, 10 options (so random guessing gives 10%)Accuracy, usually with CoTClose to saturation at the frontier
MathGSM8K Cobbe+ 2021About 8.5k grade-school word problems (1.3k test)Exact match on the final numberSaturated. Useful only for small models and contamination studies
MATH Hendrycks+ 202112.5k competition problems, 5 difficulty levels. MATH-500 is the common subsetExact match after normalizing the answerSaturated at the frontier (above 95% on MATH-500 for reasoning models)
AIME (yearly) / FrontierMath Glazer+ 2024AIME: 30 integer-answer olympiad-qualifier problems per year. FrontierMath: unpublished research-level problemsExact match. Often avg@k or pass@1 averaged over samplesAIME 2024/2025 near saturation for top reasoning models. Each year's fresh set is used for contamination control. FrontierMath's top tiers still discriminate
ScienceGPQA (Diamond) Rein+ 2023"Google-proof" graduate-level biology, physics and chemistry multiple-choice. Diamond subset of about 198 questionsAccuracyApproaching its ceiling: frontier models above the expert-validator baseline of about 65–70%. The small n makes it noisy
CodeHumanEval Chen+ 2021164 hand-written Python function problems with unit testspass@kSaturated and widely contaminated
MBPP Austin+ 2021About 1k crowd-sourced entry-level Python problemspass@kSaturated
LiveCodeBench Jain+ 2024Problems pulled continuously from LeetCode, AtCoder and Codeforces, tagged by release datepass@1. Evaluate only on problems released after the model's cutoffActive. The date-windowing is a contamination control
SWE-bench / Verified Jimenez+ 2023Real GitHub issues from 12 Python repos. The agent edits the repo, and hidden tests decide pass/fail. Verified is a 500-task subset that humans screened% resolvedOpenAI stopped reporting Verified in Feb 2026, citing flawed tests and evidence that frontier models had memorized gold patches OpenAI 2026
SWE-bench Pro Deng+ 20251,865 longer-horizon tasks from 41 repos, with public, held-out and commercial (private) splits% resolvedActive. Now recommended by OpenAI in place of Verified
Agentsτ-bench / τ²-bench Yao+ 2024 Barres+ 2025A tool-using agent serves a simulated user (retail, airline, telecom) under domain policies. τ² adds dual control, where the user also acts on the environmentTask success checked against the final database state. pass^k = all k trials succeed (a reliability measure)Active. pass^k exposes inconsistency that pass@1 hides
Terminal-Bench 2.0 Merrill+ 202689 hard tasks in containerized terminals (sysadmin, builds, crypto, scientific porting) with human-written solutions and tests% tasks passedActive. The paper reports frontier agents below 65% at publication
OSWorld Xie+ 2024369 real desktop tasks (Ubuntu/Windows apps) for multimodal computer-use agentsExecution-based checkersActive, rising fast. Human baseline about 72%
BrowseComp Wei+ 2025Hard-to-find facts that need persistent web browsing. Answers are short and easy to verifyAccuracyActive
METR time horizon Kwa+ 2025A meta-metric: the length of task, in human-expert time, that the AI completes with 50% successLogistic fit of success against log(human time)Widely cited. Reported doubling about every 7 months since 2019
Long contextNeedle-in-a-haystack Kamradt 2023Plant one fact at depth d in a context of length L, then ask for it. Results are shown as a heatmapRetrieval accuracySaturated. Literal retrieval is too easy and says little about real long-context use
RULER Hsieh+ 2024Synthetic tasks with configurable length: multi-needle, multi-hop tracing, aggregation, QAAccuracy versus length. Defines the "effective context length"Useful. It shows most models degrade well before their advertised window
Frontier / hardHumanity's Last Exam Phan+ 20252,500 expert-written closed-ended questions, some multimodal, filtered so that frontier models failed them when the set was builtAccuracy plus calibration errorFrontier models went from single digits (Jan 2025) to roughly 40–65% by late 2026, depending on tracker and tool use
ARC-AGI-1/2 Chollet 2019 Chollet+ 2025Grid-transformation puzzles that test few-shot abstraction from "core knowledge" priorsExact-match grids, often with an efficiency (cost per task) axisARC-AGI-1 largely beaten. ARC-AGI-2's Kaggle 2025 top score was 24% on the private set Chollet+ 2026
ARC-AGI-3 ARC Prize 2026Interactive turn-based game environments where the agent must explore, infer goals and plan, with no instructionsAction-efficiency relative to humansLaunched March 2026. At release, frontier systems scored very low while humans solved all environments. Progress is fast and contested
Human preferenceChatbot Arena / LMArena Chiang+ 2024Crowd users chat with two anonymous models and voteBradley–Terry "Elo" with CIs, plus style-controlled variantsVery influential. Bias and gaming critiques in Singh+ 2025
FactualitySimpleQA Wei+ 2024About 4.3k short fact-seeking questions written adversarially against GPT-4Correct / incorrect / not attempted, which rewards abstainingUseful for measuring hallucination vs. abstention
Economic tasksGDPval Patwardhan+ 2025, MLE-bench Chan+ 2024Real deliverables from professional occupations, graded by experts (GDPval). Kaggle competitions (MLE-bench)Expert win rate vs. human professionals; medal rateNewer, expensive to grade

Notes on specific figures: the GPQA expert baseline and OSWorld human baseline are approximate values from the respective papers. HLE frontier ranges are taken from third-party trackers, which disagree with each other.

Interview angle

"Which benchmarks would you trust today and why?" Don't recite a list. Give criteria. (1) Headroom: frontier scores well below the ceiling, and the ceiling isn't capped by label noise. (2) Contamination resistance: held-out, private or continuously refreshed items, or an exact release-date cutoff. (3) Grading validity: execution or state checks rather than string match, and audited tests. (4) Enough items for the differences you care about. (5) Construct validity: does it measure the capability your product needs? Then name examples: SWE-bench Pro / Terminal-Bench for agentic coding, LiveCodeBench for algorithmic coding, HLE / FrontierMath for frontier reasoning, τ²-bench pass^k for reliability, and your own private evals above all.

Go deeper

Metrics: accuracy, pass@k, maj@k, pass^k, Elo

Accuracy and its variants

For multiple-choice or exact-answer tasks the score is just the mean of per-item 0/1 correctness. The details that matter: how you extract the answer (regex over the last "Answer: X", \boxed{} parsing, symbolic equivalence via sympy for math), how you handle refusals or unparseable outputs (count them as wrong, and report the parse-failure rate), and whether you report pass@1 averaged over several samples (often written avg@k, mean@k or pass@1 with n samples). For small sets like AIME, averaging over 8–64 samples per problem reduces sampling variance. It does not reduce the variance that comes from having only 30 problems.

pass@k and the unbiased estimator

pass@k is the probability that at least one of k independent samples is correct. It models a setting where a verifier, such as unit tests, can pick the good sample. The naive approach (draw exactly k samples, check if any pass) has high variance. Codex instead draws \(n \ge k\) samples per problem, counts the \(c\) correct ones, and uses the unbiased estimator Chen+ 2021:

$$\widehat{\text{pass@}k} \;=\; \mathbb{E}_{\text{problems}}\left[\,1-\frac{\binom{n-c}{k}}{\binom{n}{k}}\,\right].$$

Read it as "1 minus the probability that a random size-k subset of the n samples contains no correct one". Computing the binomials directly overflows, so use the stable product form \(1-\prod_{i=n-c+1}^{n}\left(1-\frac{k}{i}\right)\) (and return 1 if \(n-c \lt k\)). Codex used n = 200 for k up to 100.

Why not just compute \(1-(1-\hat p)^k\) with \(\hat p = c/n\)? Because it's biased: \(1-(1-p)^k\) is concave in \(p\), so plugging in a noisy \(\hat p\) systematically underestimates the true value (Jensen's inequality). The combinatorial form is exactly unbiased for \(k \le n\).

00.50.9 124 1664256 k (samples, log scale) pass@k RL-tuned: higher pass@1, plateaus Base/diverse: lower pass@1, keeps rising curves cross
Illustrative shapes, not real data. pass@k grows with k at a rate set by sample diversity. Sharpening the distribution (RL, low temperature) raises pass@1 but can flatten the curve, so at large k a more diverse model can overtake it. This crossover has been reported in debates over whether RLVR adds new capability or mainly re-weights existing samples.

Practical notes on pass@k: temperature matters (higher T helps large-k, hurts k=1, so the Codex paper tuned T per k); pass@k with large k is only meaningful when a verifier exists at deployment; and it's increasingly used as a probe for "latent capability": if pass@256 is high but pass@1 is low, RL may be able to surface the capability. Whether RL with verifiable rewards expands pass@k at large k or only sharpens pass@1 is an active and contested research question as of 2025–2026.

pass^k (reliability)

τ-bench introduced pass^k: the probability that all k independent trials succeed. With per-task success rate \(p_t\) it is \(\mathbb{E}_t[p_t^{\,k}]\), estimated unbiasedly as \(\binom{c}{k}/\binom{n}{k}\) Yao+ 2024. A model with 60% pass@1 that's consistent per task (some tasks 100%, others 0%) keeps a high pass^8. One that succeeds on every task 60% of the time collapses to 0.6⁸ ≈ 1.7%. Customer-facing agents need pass^k, not pass@k.

maj@k (self-consistency)

maj@k samples k reasoning chains, extracts each final answer, and outputs the majority vote Wang+ 2022. Unlike pass@k it produces one answer with no oracle, so it's a fair deployable comparison. It needs answers that can be compared for equality (numbers, choices), and it plateaus once the most common wrong answer dominates. Variants weight the votes by a reward model or by confidence (best-of-N with an RM). When a lab reports "maj@64", it is spending 64× inference compute, so compare at matched compute.

MetricQuestion it answersNeeds oracle at deploy?Typical use
pass@1 (avg over n)Expected single-shot accuracyNoDefault headline number
pass@kCan the model produce a correct answer at all within k tries?Yes (tests, verifier)Code with unit tests, latent-capability probes
pass^kDoes it succeed every time?NoAgent reliability
maj@k / cons@kAccuracy of the voted answer from k samplesNoMath and multiple-choice with test-time compute
best-of-N (RM)Accuracy when a learned verifier picks one of NLearned RMTest-time scaling studies

Elo and Bradley–Terry for arenas

Arena platforms collect pairwise votes. The Bradley–Terry (BT) model assigns each model a strength \(\beta_m\) and assumes

$$P(A \succ B) = \sigma(\beta_A-\beta_B) = \frac{1}{1+e^{-(\beta_A-\beta_B)}}.$$

The strengths are fit by maximum likelihood, which is just logistic regression where each battle is a row with +1 for A's column and −1 for B's. They're reported on an Elo-like scale, \(R = 400\,\beta/\ln 10 + \text{const}\), so that \(P(A\succ B)=1/(1+10^{(R_B-R_A)/400})\). A 100-point gap is about a 64% win rate, and 200 points about 76%. Chatbot Arena moved from online Elo updates, which depend on the order of games, to batch BT fitting with bootstrap confidence intervals Chiang+ 2024. Ties are handled by counting half wins or with a Rao–Kupper/Davidson extension.

Known issues with arenas:

Interview angle

"Model A has Elo 1300, B has 1280. Is A better?" The strong answer: a 20-point gap is about a 53% win rate. Check whether the bootstrap CIs overlap (with tens of thousands of votes they're often ±5–10 points, but much wider for new models). Ask whether the score is style-controlled, which category it's in, and whether it would hold on your traffic. Then run your own blind pairwise eval on in-domain prompts.

Go deeper

Scoring protocols: shots, prompts, log-likelihood vs generation

Zero-shot vs few-shot

Few-shot prompting puts k solved examples in the context. For base models it's essential: it tells the model the format and the task, and 5-shot MMLU was the standard. For chat/instruct models, zero-shot with an instruction (often plus "think step by step" or native reasoning) is now the norm. Few-shot exemplars can even hurt reasoning models by constraining their reasoning format. Results are not comparable across shot settings, and the choice and order of exemplars add variance on their own.

Prompt sensitivity

Meaning-preserving formatting changes (separators, casing, spacing, "Answer:" vs "A:") swung LLaMA-2-13B few-shot accuracy by up to 76 points on some tasks, and the best format barely correlated across models Sclar+ 2023. Frontier instruction-tuned models are much more robust, but sensitivity of a few points is still common, which is the same size as many claimed improvements. GSM-Symbolic showed that changing names and numbers in GSM8K templates, or adding irrelevant clauses, reduces accuracy and increases variance, a sign of pattern-matching rather than robust reasoning Mirzadeh+ 2024. Good practice is to report the mean and spread over several prompt variants, and to use the model's own chat template.

Log-likelihood (cloze/MC) vs generative scoring

Log-likelihood scoring

Compute \(\log p(\text{choice}_j \mid \text{prompt})\) for each option and pick the argmax. You can score the full answer text ("cloze") or just the letter ("MC format").

  • Cheap: one forward pass, no sampling, deterministic
  • Works on small or base models that can't follow output formats yet, so it's ideal for pretraining ablations
  • Needs length normalization (per token, per character, or relative to an unconditional prior), and the choice changes rankings
  • Can't evaluate reasoning chains, and isn't possible for closed APIs without logprobs

Generative scoring

Sample (or greedy-decode) a full response, extract the answer, and compare.

  • Matches how the model is actually used, and allows CoT and reasoning
  • Brittle answer extraction: format failures get counted as knowledge failures
  • Sampling adds variance, so use multiple samples
  • Much more expensive, especially with long reasoning traces

The same model can rank differently under the two schemes. Small models often look much better under cloze log-likelihood, because the letter-MC format is itself a capability that emerges only at scale. The lm-eval-harness authors document these choices in detail Biderman+ 2024.

Common mistake

Comparing your number against a number in someone else's paper. Unless the harness version, prompt, shots, chat template, sampling settings and answer extraction all match, differences of several points are expected. Re-run the baselines yourself in the same harness.

Go deeper

Contamination

How it happens

Detection methods

MethodHow it worksNeedsLimits
N-gram / substring overlapSearch the training corpus for long n-gram matches with test items. GPT-4's report used 50-character substring matching OpenAI 2023, and Llama 3 reports similar token n-gram analyses Grattafiori+ 2024Training data accessMisses paraphrases and translations. Thresholds are arbitrary
Membership inference (Min-K% Prob)Average log-prob of the k% least likely tokens in a text. Seen text has fewer very-surprising tokens Shi+ 2023Logprobs onlyNoisy per example. Better for aggregate evidence
Exchangeability / order testIf the model assigns higher likelihood to the benchmark's canonical item order than to shuffled orders, it probably saw the file. This gives a statistical guarantee Oren+ 2023LogprobsOnly detects verbatim, ordered inclusion
Completion / regurgitation probesPrompt with the first half of an item and check whether the model completes it verbatim (or reproduces the gold patch)Black-box samplingAbsence isn't proof of cleanliness
Fresh mirror setBuild new items with the same distribution and difficulty, then compare. GSM1k found drops of up to 8% for some model families, with frontier models showing little overfitting Zhang+ 2024Annotation budgetExpensive, and hard to match difficulty exactly
Perturbation testsChange numbers, names or option order, or rephrase. Accuracy that drops on the perturbed items suggests memorizationTemplatesPerturbations can change difficulty
Time-split ("live") evalsOnly score items created after the training cutoff (LiveCodeBench, LiveBench White+ 2024, each new AIME)Continuous item supplyCutoff dates are fuzzy, and the items get stale too

Mitigations as a benchmark designer

Keep a private held-out split and evaluate through a server (SWE-bench Pro's held-out and commercial splits, ARC-AGI private sets, FrontierMath). Add canary strings (a unique GUID in every file so that crawlers and labs can filter it, as BIG-bench did), keep data gated behind terms of use, refresh items continuously, and prefer tasks where memorizing the answer doesn't help (interactive environments, procedurally generated tasks).

Interview angle

"Your model jumps 10 points on a public benchmark after a data refresh. What do you check?" Look for n-gram overlap between the new data and the test set. Compare against a held-out mirror or perturbed set. Check per-item: are the gains concentrated on items that appear on the web? Run completion probes. Check whether gains transfer to fresh, time-split items. Strong candidates also mention adaptive overfitting (you chose this checkpoint because of this benchmark) and propose a quarantined final test set that is looked at rarely.

Go deeper

Saturation, the benchmark lifecycle and Goodhart

"When a measure becomes a target, it ceases to be a good measure" (Goodhart's law, in Strathern's popular phrasing). Benchmarks follow a predictable arc:

introduce adopt optimize-against saturate / die ┌──────────┐ ┌──────────────┐ ┌───────────────────────┐ ┌────────────────────────┐ │ hard, │──▶ │ cited in │──▶ │ labs tune data mixes, │──▶ │ near ceiling; residual │ │ low SOTA │ │ model cards, │ │ RL envs, prompts; │ │ errors = label noise; │ │ (~10-30%)│ │ leaderboards │ │ test leaks to web │ │ gains don't transfer │ └──────────┘ └──────────────┘ └───────────────────────┘ └───────────┬────────────┘ ▲ │ └──────────── harder successor: MMLU→MMLU-Pro→HLE, GSM8K→MATH→AIME→FrontierMath, HumanEval→LiveCodeBench, SWE-bench Verified→SWE-bench Pro, ARC-1→2→3

Signs that a benchmark is saturated: (1) frontier scores are within a few points of the label-noise ceiling (for example, MMLU's estimated 6.5% error rate means about 93% is a soft ceiling); (2) score gains no longer predict gains on fresh or adjacent tasks; (3) top-model differences fall inside the confidence intervals; (4) evidence of memorization appears. The time from release to saturation has shrunk from years to months for some recent benchmarks, which pushes the field toward private, live and agentic evals that are costly to game.

Goodhart failure modes specific to LLMs:

Intuition

A benchmark is a cheap proxy for a capability. Its value lies in the correlation between proxy and capability across models that weren't optimized for it. Optimization pressure finds the directions where proxy and capability come apart, so the correlation always decays. The only defenses are to rotate proxies and keep some private.

LLM-as-judge: biases and mitigations

For open-ended outputs, a strong model grades either a single response (pointwise, for example on a 1–10 rubric) or a pair (pairwise preference). MT-Bench showed that GPT-4 judgments agreed with human experts roughly as often as humans agreed with each other (above 80% on non-tie pairs), which legitimized the approach. The same paper catalogued its biases Zheng+ 2023.

BiasWhat happensMitigation
Position biasIn pairwise comparisons the judge favours the first (sometimes the second) slot. Swapping the order can flip verdicts on a large fraction of pairs Wang+ 2023Evaluate both orders. Count a win only if it's consistent, else a tie, or average the probabilities
Verbosity / length biasLonger, more detailed-looking answers are preferred even when they aren't betterRubrics that penalize padding. Length-controlled win rates (a regression adjustment, as in Length-Controlled AlpacaEval Dubois+ 2024). Report length alongside the score
Self-preferenceJudges rate their own (or same-family) outputs higher. This correlates with the judge's ability to recognize its own text Panickssery+ 2024Use a judge from a different family, or a jury of diverse judges (PoLL Verga+ 2024)
Style / authority biasMarkdown, confident tone and citations (even fake ones) sway the judgeNormalize formatting before judging. Check facts separately
Limited capabilityThe judge can't verify math or code it can't solve itselfGive the judge reference answers, execution results or tests. Use a stronger judge than the model being graded
Score miscalibrationPointwise 1–10 scores cluster and drift across judge versionsPairwise or binary rubric items, anchored examples, and a pinned judge model version

Process for a trustworthy judge: write a decomposed rubric (several binary checks rather than one holistic score); label a few hundred items with humans; measure judge–human agreement (accuracy, Cohen's κ) and compare it to human–human agreement; iterate the prompt; then freeze the judge version and re-validate whenever you change it. Use the judge for relative comparisons, and don't use the same judge as an RL reward and as the final evaluator.

Interview angle

They'll ask "how do you know your judge is right?" The answer is measured agreement with human labels on a stratified sample, compared to the human ceiling, plus checks for each known bias: swap order, length-matched pairs, and a different-family judge. Bonus points for spotting the RL trap: a policy optimized against a judge will find that judge's blind spots, so hold out a different judge (or humans) for evaluation.

Statistical rigor: error bars on evals

Treat an eval as an experiment: the questions are a sample from a larger super-population of possible questions, and the score is an estimate with sampling error Miller 2024 Anthropic 2024.

Standard error of the mean

For n items with per-item scores \(s_i\) (0/1 for accuracy), \(\mathrm{SE} = \sqrt{\widehat{\mathrm{Var}}(s)/n}\), which is \(\sqrt{p(1-p)/n}\) for accuracy. The 95% CI is about \(\pm 1.96\,\mathrm{SE}\).

n (items)ExampleSE at p = 0.595% CI half-widthat p = 0.9
30one AIME year9.1 pts±17.9 pts±10.7 pts
198GPQA Diamond3.6 pts±7.0 pts±4.2 pts
500SWE-bench Verified, MATH-5002.2 pts±4.4 pts±2.6 pts
2,500HLE1.0 pts±2.0 pts±1.2 pts
14,000MMLU0.42 pts±0.83 pts±0.50 pts

To halve the error bar you need 4× the items. A useful rule: the minimum detectable difference between two independent evaluations at 80% power and α = 0.05 is about \(2.8\sqrt{2}\,\mathrm{SE}\), roughly 4 SE. So on a 500-item benchmark near 50%, you need a gap of about 9 points to reliably detect it with unpaired analysis.

Paired comparisons (the big win)

When two models answer the same questions, analyse per-question differences \(d_i = s_{A,i}-s_{B,i}\):

$$\mathrm{Var}(\bar d) = \frac{\mathrm{Var}(s_A)+\mathrm{Var}(s_B)-2\,\mathrm{Cov}(s_A,s_B)}{n}.$$

Because models tend to find the same questions easy or hard, the covariance is usually substantial and positive, so the paired SE can be much smaller than the unpaired one. Paired analysis is "free variance reduction" Miller 2024. Use a paired t-test on \(d_i\), McNemar's test on the discordant pairs (A right/B wrong vs. A wrong/B right), or a paired bootstrap.

Other variance sources

Interview angle

Classic back-of-envelope question: "New model scores 62% vs 59% on a 500-question benchmark. Ship it?" SE per model is about 2.2 points, so an unpaired 95% CI on the difference is about ±6 points: not significant on its own. With paired analysis and high correlation it might be. Ask for the per-item results, run McNemar or a paired bootstrap, check the discordant counts, look at other benchmarks for the same direction, and consider re-sampling to reduce generation noise.

Go deeper

Evaluating reasoning and test-time compute

Reasoning models (o1-style and successors) turn inference compute into accuracy, both through longer chains of thought and through parallel sampling. A single accuracy number hides this trade. Snell et al. showed that compute-optimal test-time strategies, adapted to problem difficulty, can beat a much larger model at matched FLOPs on some problems Snell+ 2024. s1 showed that "budget forcing" (suppressing the end-of-thinking token, or appending "Wait") lets you sweep thinking length and trace a curve Muennighoff+ 2025.

How to evaluate properly:

reasoning tokens per problem (log scale) accuracy 1k4k16k64k Model A: efficient, plateaus Model B: needs budget, higher ceiling
Illustrative. The "better" model depends on your token budget. Report the whole curve, or at least accuracy at two or three matched budgets.
May be out of date

Conventions for reporting reasoning effort (effort levels, thinking budgets, adaptive thinking) changed several times in 2025–2026 and differ by vendor. Check how each vendor defines its effort settings before comparing.

Safety and dangerous-capability evals, red teaming

Frontier labs run dangerous capability evaluations before release under their scaling and preparedness policies. The main domains are CBRN uplift (chemical, biological, radiological, nuclear), offensive cyber, autonomous replication and AI R&D automation, and persuasion or deception. Google DeepMind's 2024 paper is a good template for how the domains are structured Phuong+ 2024. Typical forms:

TypeExampleWhat makes it different
Knowledge proxiesWMDP: about 3.7k multiple-choice questions on hazardous bio/chem/cyber knowledge, also used to measure unlearning Li+ 2024Cheap, but knowledge isn't the same as uplift
Agentic task suitesCapture-the-flag cyber challenges, autonomous ML-engineering tasks (MLE-bench), self-proliferation tasks, METR-style long tasksMeasure end-to-end ability with tools. Expensive and sensitive to scaffolding
Human uplift studiesRandomized trials: experts or novices with vs. without model access on proxy tasksThe gold standard for "uplift", but slow and costly
Red teamingHuman experts and automated attackers try to elicit harmful outputs or bypass safeguardsAdversarial and open-ended, so it finds unknown unknowns
Propensity / alignment evalsDoes the model deceive, sandbag, seek power or reward-hack when given the chance?Measures disposition, not just capability. Hard to rule out eval-awareness

Elicitation is the core methodological issue. A capability eval is an upper-bound claim only if you tried hard to elicit the capability: best scaffolding, tools, fine-tuning to remove refusals, and many attempts. Under-elicitation leads to false "safe" conclusions. Sandbagging (a model strategically underperforming) and evaluation awareness (models recognizing tests) are recent concerns that make propensity evals harder.

Red teaming: manual red teams of domain experts, plus automated red teaming where an attacker LM generates test cases and a classifier scores them Perez+ 2022. Anthropic's scaling study of human red teaming found RLHF models got harder to red-team as they scaled, and released the attack dataset Ganguli+ 2022. Jailbreak robustness is measured as attack success rate (ASR) under a fixed attack budget.

May be out of date

Lab frameworks (Anthropic RSP / ASL levels, OpenAI Preparedness Framework, Google DeepMind Frontier Safety Framework) and government evaluation institutes (UK AISI, US CAISI) revised thresholds and processes repeatedly in 2025–2026. Check the current versions directly.

Interview angle

"How is a dangerous-capability eval different from a normal benchmark?" (1) You want an upper bound, so elicitation effort is part of the method. (2) The quantity that matters is uplift relative to baseline resources (internet, experts), not raw accuracy. (3) Thresholds trigger actions, so false negatives cost far more than false positives. (4) Test sets must stay private, and some are hazardous. (5) Results have to hold up against sandbagging.

Go deeper

Eval harnesses

HarnessWhat it isStrengthsNotes
lm-evaluation-harness (EleutherAI)Hundreds of tasks defined in YAML. Log-likelihood and generative modes. HF, vLLM and API backendsThe standard for base-model and open-model evaluation, and the backend of the original Open LLM Leaderboard. Reproducibility focus Biderman+ 2024Agentic and multi-turn support is limited
HELM (Stanford CRFM)Holistic evaluation: many scenarios × metrics (accuracy, calibration, robustness, fairness, toxicity, efficiency) Liang+ 2022Standardized, transparent comparisons, with domain-specific leaderboardsMore of a benchmark framework than a dev-loop tool
Inspect (UK AISI)Python framework: datasets → solvers (prompting, tool use, agents) → scorers (including model-graded), with sandboxed execution (Docker)Strong for agentic and safety evals. Large library of community evalsWidely adopted by safety orgs
simple-evals (OpenAI)Minimal reference implementations: MMLU, MATH, GPQA, DROP, MGSM, HumanEval, SimpleQA, BrowseComp, HealthBenchZero-shot, chat-style prompts, transparentThe README states it is no longer updated for new models as of July 2025
Agent harnesses (SWE-bench harness, Harbor for Terminal-Bench, OSWorld env)Containerized environments plus task-specific gradersExecution-based gradingThe scaffold is part of the result. Report it

Whatever harness you use, log every raw transcript, version everything (model, harness commit, prompts, judge model), and store per-item scores so you can run paired analysis later. Most "mysterious" eval regressions turn out to be template, tokenizer or extraction bugs, and you find them by reading transcripts.

Part 2 · Interpretability: why it matters

Evals tell you what a model does on the inputs you thought to test. Interpretability tries to tell you how it does it, which gives you predictions about inputs you didn't test. The main motivations:

A useful axis: top-down / representation-level methods (probing, steering, representation engineering) find directions that correlate with concepts and test interventions. Bottom-up / mechanistic methods (circuits, SAEs, attribution graphs) try to reverse-engineer the computation into human-understandable parts. Another axis is correlational vs causal: probes and attention maps show what is present, while patching and ablation show what is used.

Interview angle

For an AI engineer role (not a research scientist role), interviewers mostly probe whether you understand the core ideas (residual stream, superposition, SAEs, probing vs patching) and their limitations, and whether you can connect them to practice: probes as cheap classifiers on activations (for example, harmful-intent detection), steering vectors for behaviour control, and why CoT shouldn't be treated as an explanation.

Probing (linear probes)

A probe is a small classifier, usually linear (logistic regression), trained on frozen hidden activations \(h_\ell(x)\) at layer \(\ell\) to predict a property \(y\): part of speech, sentiment, truthfulness, "is this prompt harmful", board state in Othello-GPT. If a linear probe gets high held-out accuracy, the information is linearly decodable at that layer Alain & Bengio 2016. Sweeping layers gives a profile of where a feature becomes available.

Caveats

Unsupervised variants: Contrast-Consistent Search (CCS) finds a direction where a statement and its negation get complementary probabilities, with no labels, as an attempt to find "what the model believes" Burns+ 2022. Later work questioned whether CCS finds truth specifically rather than other salient features. The broader linear representation hypothesis (concepts are directions) has both theoretical framing Park+ 2023 and known counterexamples, such as circular features for days of the week and other nonlinear manifolds.

Intuition

Probes are a strong, cheap practical baseline. In GDM's 2025 work on detecting harmful intent, linear probes on raw activations beat SAE-based classifiers out of distribution GDM 2025. If you need a classifier over model internals, start with a linear probe.

The logit lens and tuned lens

Transformers keep a residual stream \(h_\ell\) that every layer adds to. The final prediction is \(\mathrm{softmax}(W_U\,\mathrm{LN}(h_L))\). The logit lens applies the unembedding to an intermediate residual: \(\mathrm{softmax}(W_U\,\mathrm{LN}(h_\ell))\). This lets you watch the prediction form layer by layer: early layers decode to junk or the input token, then the right answer "rises" in the middle and late layers nostalgebraist 2020.

It works because the residual stream is roughly aligned to a shared basis across layers, but it's brittle: in some models (early layers, or models whose basis drifts) the decoded distributions are meaningless. The tuned lens trains a small affine "translator" per layer, \(\mathrm{softmax}(W_U\,\mathrm{LN}(A_\ell h_\ell + b_\ell))\), to minimize KL to the final output. It's more reliable, and the trajectory of predictions across depth can even help detect prompt injection or anomalous inputs Belrose+ 2023.

Common mistake

Reading the logit lens as "what layer ℓ thinks the answer is". Intermediate layers also store information intended for later layers, in directions the unembedding doesn't read. An absence in the logit lens doesn't mean an absence in the representation.

Attention-pattern analysis and its limits

Attention weights \(A = \mathrm{softmax}(QK^\top/\sqrt{d})\) are easy to visualize. Head-level patterns are real and sometimes interpretable: previous-token heads, induction heads, heads that attend to a syntactic head or to the first token ("attention sinks"). Tools like BertViz and CircuitsVis made this popular.

Why attention isn't explanation:

The mechanistic replacement is to separate each head into a QK circuit (\(W_E^\top W_Q^\top W_K W_E\): which tokens attend to which) and an OV circuit (\(W_U W_O W_V W_E\): what gets written when attended), and to test claims causally Elhage+ 2021. Anthropic's 2025 work extends attribution to explaining why attention patterns form, in terms of feature interactions.

Mechanistic interpretability: residual stream, features, superposition, circuits

The residual-stream view

"A Mathematical Framework for Transformer Circuits" reframed the transformer as a communication channel. The residual stream is a shared memory of width \(d_\text{model}\). Each attention head and MLP reads from it through a linear projection, computes, and adds its output back Elhage+ 2021. Consequences:

Features and superposition

A feature is a property of the input the network represents, ideally as a direction in activation space. If each neuron represented one feature, neurons would be monosemantic. In practice many are polysemantic, firing for unrelated things (academic citations, Korean text and HTTP requests). Superposition explains this: when features are sparse (rarely active at the same time), a network can represent many more features than it has dimensions by assigning them nearly orthogonal directions. It accepts a little interference in exchange for capacity Elhage+ 2022. In high dimensions you can fit exponentially many almost-orthogonal vectors (by the Johnson–Lindenstrauss intuition), so superposition is cheap. The toy models show phase changes: as sparsity rises, features move from "not represented" to "dedicated dimension" to "superposed in geometric arrangements" such as antipodal pairs and pentagons.

Intuition

Superposition is compression. The model is a lossy code for a huge number of rare features, packed into a few thousand dimensions. Neurons are the wrong basis for reading that code. You need to recover the overcomplete feature basis, and that's the motivation for sparse dictionary learning.

Circuits and induction heads

A circuit is a subgraph of features and weights that implements a behaviour. The canonical example is the induction head, a two-head circuit that does in-context copying: on a sequence … [A][B] … [A] it predicts [B] Olsson+ 2022.

tokens: ... Harry Potter ... ... Harry → predict "Potter" pos i pos i+1 pos j Layer 1 · previous-token head: at pos i+1 writes "previous token was Harry" into the residual Layer 2 · induction head: at pos j, query = "current token is Harry" key (via K-composition) = "my previous token was Harry" → matches pos i+1 OV copies the token at i+1 ("Potter") → boosts its logit

Olsson et al. found that induction heads form in a sudden phase change early in training. It shows up as a bump in the loss curve and coincides with a jump in in-context learning ability (loss on later tokens falling relative to earlier ones), and the paper argues that induction heads may account for much of in-context learning in larger models. Other well-studied circuits include the indirect object identification (IOI) circuit in GPT-2 small, which has name-mover, S-inhibition and duplicate-token heads Wang+ 2022, and automated discovery methods such as ACDC Conmy+ 2023.

Activation patching and causal tracing

The workhorse causal method. Run the model on a clean input (for example "The Eiffel Tower is in" → Paris) and a corrupted input (a different subject, or noise added to the subject embeddings). Then, one at a time, copy an activation (a head output, MLP output or residual at some layer and position) from one run into the other, and measure how much of the clean behaviour is restored, for example through the logit difference.

The methodological choices (corruption type, metric, which positions) noticeably change the conclusions. Gaussian-noise corruption can push the model off-distribution, and logit difference is generally preferred over raw probability Zhang & Nanda 2023 Heimersheim & Nanda 2024. Watch for self-repair (backup heads compensate when one is ablated), which makes necessity tests understate importance.

Interview angle

"How would you find which part of the model stores fact X?" Describe causal tracing: build a clean/corrupted pair, sweep patches over (layer, position), find the restoration peak, then confirm with targeted ablations and an edit (ROME). Mention the pitfalls: off-distribution corruption, self-repair, and the fact that localization doesn't guarantee that editing there will work well (see the model-editing section).

Go deeper

Sparse autoencoders and dictionary learning

If activations \(x \in \mathbb{R}^{d}\) are sparse linear combinations of many feature directions, \(x \approx \sum_i f_i\, d_i\) with most \(f_i = 0\) and \(m \gg d\) directions, then recovering \(\{d_i\}\) is a sparse dictionary learning problem. A sparse autoencoder (SAE) learns it with a one-layer encoder and a linear decoder Cunningham+ 2023 Bricken+ 2023:

$$f(x) = \mathrm{ReLU}\big(W_e (x - b_d) + b_e\big) \in \mathbb{R}^{m}, \qquad \hat x = W_d\, f(x) + b_d,$$ $$\mathcal{L} = \underbrace{\lVert x - \hat x\rVert_2^2}_{\text{reconstruction}} \;+\; \lambda \underbrace{\sum_i f_i(x)\,\lVert W_{d,:,i}\rVert_2}_{\text{sparsity (L1)}}.$$

Each column of \(W_d\) is a feature direction. The expansion factor \(m/d\) typically ranges from 8× to 256× or more. Each latent's activations are inspected (top-activating examples, logit effects) and auto-labelled by an LLM.

residual x d = 4096 dense, polysemantic encoder W_e, ReLU (or TopK / JumpReLU) features f(x) "Golden Gate Bridge" "code: Python list" "sarcasm" m = 64k–16M wide; only ~10–100 active (L0) decoder W_d x̂ ≈ x
An SAE expands a dense, polysemantic activation into a very wide, sparse code whose latents are (ideally) monosemantic features, then reconstructs the input linearly. Loss = ‖x − x̂‖² + λ·sparsity. Feature labels are illustrative. Sizes are typical ranges from published work.

Key results and variants

How SAEs are evaluated

MetricMeaningTrade-off
L0Average number of active latents per tokenLower is sparser and more interpretable, but reconstructs worse
Fraction of variance explained / MSEReconstruction qualityCan look high while missing what matters for behaviour
Δ loss (spliced-in CE)Replace activations with \(\hat x\) and measure the increase in LM lossThe most behaviour-relevant fidelity metric. Splicing typically still costs noticeable loss
Interpretability scoresAuto-interp: an LLM explains a latent, then predicts its activations from the explanationCheap at scale, but explanations can be superficial
Dead / dense latentsLatents that never fire, or fire on everythingA training-health indicator (resampling, TopK and initialization tricks help)
Downstream tasksProbing, steering, unlearning or spurious-correlation removal using SAE features vs. baselinesIncreasingly seen as the real test, and results are mixed

Feature splitting and other limitations

May be out of date

The field's view of SAEs shifted noticeably between 2024 (high enthusiasm) and 2025–2026 (more skepticism, more emphasis on downstream benchmarks, and newer architectures such as transcoders, crosscoders, sparse mixtures of linear transforms and turn-averaged SAEs). Treat "SAEs solve superposition" as contested, not established.

Interview angle

"Explain an SAE in two minutes." Superposition means neurons are polysemantic, so learn an overcomplete dictionary. Encoder with ReLU/TopK, linear decoder, loss = MSE + sparsity penalty. Features are columns of the decoder. Evaluate by L0, spliced-in Δloss and interpretability. Limitations: feature splitting, dark matter, mixed downstream utility (linear probes often win for known concepts). Mentioning the L1 shrinkage problem and the TopK/JumpReLU fix signals real depth.

Go deeper

Transcoders, attribution graphs and circuit tracing

SAEs decompose a representation. To trace computation you want interpretable pieces connected by interpretable interactions. Transcoders do this: instead of reconstructing an MLP's input, a transcoder is trained to imitate the MLP, mapping its input to its output through a sparse latent layer. You can then substitute it for the MLP, and feature-to-feature interactions become (mostly) linear and input-independent Dunefsky+ 2024.

Anthropic's circuit tracing (2025) scaled this up Ameisen+ 2025:

  1. Train a cross-layer transcoder (CLT): each layer's features read from that layer's residual stream and can write to all subsequent MLP outputs, which collapses multi-layer MLP computation into fewer features.
  2. Build a replacement model by swapping all MLPs for the CLT. For a given prompt, freeze attention patterns and normalization, and add error nodes for whatever the CLT fails to reconstruct, so that the local replacement exactly reproduces the model's output on that prompt.
  3. Because the replacement is linear given the frozen attention, compute an attribution graph: nodes are active features (plus input tokens, error nodes and output logits), and edges are direct linear contributions. Prune it to the most influential paths and group features into "supernodes" by hand.
  4. Validate hypotheses with interventions in the real model: suppress or inject features and check that downstream features and outputs change as the graph predicts.

Findings in Claude 3.5 Haiku Lindsey+ 2025 included genuine multi-hop reasoning inside the forward pass ("capital of the state containing Dallas" goes through an intermediate "Texas" feature), planning in poetry (rhyme candidates chosen before the line is written), language-independent concept features with language-specific input and output circuitry, a default "can't answer" circuit that known-entity features inhibit (relevant to hallucination), and cases where the model's verbal explanation of its arithmetic didn't match the internal mechanism. The authors noted that the method gave satisfying insight on only roughly a quarter of the prompts they tried, and that it doesn't explain attention-pattern formation (QK) by default. An open-source circuit-tracer library with Neuronpedia integration lets you generate attribution graphs for open-weight models.

Recent developments (2025–2026), verify current state

Anthropic's Transformer Circuits thread kept publishing through 2026. Topics include tracing attention computation through feature interactions (2025), crosscoder model diffing, sparse mixtures of linear transforms as a transcoder alternative, emergent introspective awareness (2025), "Natural Language Autoencoders" that train Claude to verbalize its activations (2026), and a "global workspace" of verbalizable representations (2026) Transformer Circuits. These are recent and less independently replicated than the 2021–2024 core results. Read them as active research, not settled consensus.

Steering vectors and representation engineering

If concepts are (approximately) linear directions, you can add a direction to the residual stream at inference to change behaviour, with no weight updates. Activation addition (ActAdd) takes the difference of activations on a contrast pair of prompts ("Love" − "Hate") and adds a scaled version at some layer for all positions Turner+ 2023. Contrastive Activation Addition (CAA) averages the difference over many paired examples of a behaviour (sycophantic vs. not) to get a more robust vector Panickssery+ 2023:

$$v_\ell = \frac{1}{|D|}\sum_{(p^+,\,p^-) \in D} \big(h_\ell(p^+) - h_\ell(p^-)\big), \qquad h_\ell \leftarrow h_\ell + \alpha\, v_\ell .$$

Representation Engineering (RepE) generalizes this into "reading" (find concept directions with PCA on contrastive activations, use them as monitors) and "control" (add or subtract them) for honesty, harmlessness, emotion and more Zou+ 2023. A striking example: refusal in many chat models is mediated by a single direction. Ablating it (projecting it out of weights or activations) largely removes refusals, and adding it makes the model refuse harmless requests Arditi+ 2024. This shows both the power of the method and why open-weight safety tuning is shallow. Persona vectors extend the idea to monitoring and controlling traits such as sycophancy or "evil" across fine-tuning, including predicting which training data will shift a trait Chen+ 2025.

ApproachMechanismProsCons
PromptingInstruction in contextTrivial, no access to internalsCan be overridden, uses context
Steering vectors (CAA/RepE)Add a direction to the residual at inferenceCheap, reversible, continuous strength α, few examplesNeeds activation access. Large α hurts fluency and capability. Layer and α need tuning. Effects can be inconsistent across inputs
SAE feature clampingSet specific SAE latents to high or low valuesInterpretable knobNeeds a good SAE. Off-target effects
Fine-tuning (SFT/RL/LoRA)Change weightsRobust, generalCostly, less targeted, can regress other skills
Interview angle

"How would you reduce sycophancy without retraining?" Build a contrast dataset (sycophantic vs. honest completions on the same prompts), compute a CAA vector at a middle layer, subtract it at inference with α tuned on a validation set, and measure both the target behaviour and general capability (an MMLU-style regression check). Mention monitoring as an alternative use: project activations onto the vector to flag sycophantic generations.

Model editing: ROME and MEMIT (briefly)

Model editing changes a specific fact ("The Eiffel Tower is in Rome") without retraining. ROME uses causal tracing to find the mid-layer MLP that recalls facts at the last subject token. It treats that MLP's output projection as a linear associative memory \(W k \approx v\) (key = subject representation, value = fact), and applies a closed-form rank-one update so the subject's key maps to a new value while other keys are disturbed as little as possible Meng+ 2022. MEMIT spreads updates across several layers to insert thousands of facts at once Meng+ 2022.

Evaluation uses efficacy (the edit holds), generalization (paraphrases work), and specificity (neighbouring facts don't change). Known limitations: poor "ripple effects" (editing a person's birthplace doesn't update facts that follow from it), degradation after many sequential edits, and evidence that where causal tracing localizes a fact doesn't predict where editing works best. In practice, retrieval (RAG) or fine-tuning is usually preferred over weight editing for knowledge updates. Editing matters more as an interpretability probe than as a product tool.

Chain-of-thought faithfulness

A visible reasoning trace is tempting to read as an explanation. Faithfulness asks whether the CoT reflects the factors that actually drove the answer. The evidence says it often doesn't:

But CoT is still useful for monitoring. When a task is hard enough that the model needs to reason in tokens, the reasoning has to pass through the visible trace, so monitors can catch misbehaviour there. A multi-lab position paper argues that this monitorability is "a new and fragile opportunity", which training pressure on the CoT (optimizing it to look good) or architectures that reason in latent space could destroy Korbak+ 2025.

Common mistake

Treating CoT as an audit log ("the model explained why it chose this, so it's safe"), or penalizing "bad thoughts" in the CoT during RL. The second teaches the model to hide its reasoning rather than to stop the behaviour, and it degrades the CoT's monitoring value.

Interview angle

"Can we trust a reasoning model's chain of thought?" A strong answer distinguishes faithfulness (does the CoT reflect the true causes? often not, per Turpin and Chen) from monitorability (is the CoT useful for detecting problems? often yes, when reasoning is necessary). It names the tests (bias injection, hint verbalization, CoT truncation and corruption) and the policy implication: don't optimize the CoT directly, and combine CoT monitoring with output checks and interpretability-based probes.

What is established vs. recent or contested

Well establishedRecent / contested (2025–2026)
The unbiased pass@k estimator. BT/Elo for pairwise ranking. SE formulas and paired analysis. Prompt-format sensitivity. Contamination exists and inflates public benchmarks. LLM-judge position, length and self-preference biases.Which coding and agent benchmarks are still "valid" (SWE-bench Verified retired, Pro and Terminal-Bench rising). How much RLVR expands vs. sharpens capability (pass@k at large k). Arena integrity. Comparability of reasoning-effort settings. Eval awareness and sandbagging.
The residual-stream framework. Superposition in toy models. Induction heads and their link to ICL. Activation patching as a causal tool. Probes are correlational. Attention is not a full explanation. CoT is often unfaithful.How far SAEs are useful downstream. Whether dictionary features reflect the model's own ontology. How well attribution graphs scale to frontier models. Introspection and verbalization methods. How long CoT stays monitorable.

Interview question bank

Why can't you compare perplexities across two models with different tokenizers, and what do you use instead?

Perplexity is the exponentiated mean NLL per token. A tokenizer with a larger vocabulary produces fewer, longer tokens, so each prediction carries more information and per-token loss is higher even if the model is equally good per character. Normalize to a tokenizer-independent unit: bits per byte (or per character), computed as total NLL in bits divided by the number of UTF-8 bytes. You must also hold the evaluation text fixed (same documents, same preprocessing) and compare base models with base models, since post-training changes the output distribution deliberately.

Derive or explain the unbiased pass@k estimator. Why not use 1 − (1 − c/n)^k?

Generate n ≥ k samples per problem and count c correct ones. The probability that a random size-k subset contains no correct sample is C(n−c, k)/C(n, k), so pass@k = E[1 − C(n−c,k)/C(n,k)], averaged over problems. This is exactly unbiased for any k ≤ n. The plug-in 1 − (1 − p̂)^k is biased low because the function is concave in p (Jensen). For numerical stability compute 1 − ∏_{i=n−c+1}^{n}(1 − k/i), returning 1 if n − c < k. Codex used n = 200. Also tune temperature per k: higher temperatures help large-k, hurt k = 1.

A model gets pass@1 = 40% and pass@100 = 90% on a code benchmark. What does that tell you, and how would you exploit it?

The model can often produce a correct solution but can't reliably pick it, so the capability is latent and the bottleneck is selection or consistency. If unit tests or a verifier are available at deployment, generate many samples and filter (pass@k becomes a realistic deployed metric). Without one, train a verifier or reward model for best-of-N, use maj@k where outputs are comparable, or use RL with verifiable rewards to concentrate probability on correct modes, which should raise pass@1 toward pass@k. Keep in mind that RL sharpening can reduce diversity, so pass@k at very large k may fall.

What's the difference between pass@k, pass^k and maj@k? When would you report each?

pass@k: at least one of k samples correct. It needs an oracle at deployment (tests) and measures latent capability. pass^k: all k trials correct, estimated as C(c,k)/C(n,k). It measures reliability and is crucial for customer-facing agents (τ-bench). maj@k: the majority-voted answer of k samples is correct. It's deployable without an oracle and is a form of test-time compute. Report pass@1 as the headline, pass@k for code with tests, pass^k for agents, and maj@k only alongside pass@1 at matched compute.

Back-of-envelope: Model B scores 62% vs Model A's 59% on a 500-item benchmark. Is B better?

SE per model ≈ √(0.6·0.4/500) ≈ 2.2 points, so an unpaired SE of the difference is ≈ 3.1 points and the 95% CI is ≈ 3 ± 6 points: not significant. With paired data the variance subtracts 2·Cov. If the per-item correlation is high, the paired SE might be around 1.5–2 points, and a 3-point gap could become borderline significant. Run McNemar's test on the discordant items or a paired bootstrap. Also check consistency across seeds, prompts and related benchmarks, the multiple-comparisons context, and whether the benchmark's label noise makes small differences meaningless.

How many questions do you need to detect a 2-point accuracy improvement around 70% accuracy?

Unpaired: the SE of a difference is √(2·0.7·0.3/n). For 80% power at α = 0.05 you need the effect ≈ 2.8 SE, so 0.02 = 2.8·√(0.42/n), giving n ≈ 0.42·(2.8/0.02)² ≈ 8,200 questions per model. Paired analysis with correlated outcomes can cut this by a large factor (often 2–5×, depending on correlation), which is why you always use paired designs. For small benchmarks like GPQA Diamond (198 items) a 2-point difference is fundamentally undetectable.

Explain Bradley–Terry and how Elo is derived for an arena. What does a 50-point gap mean?

BT models P(A beats B) = σ(β_A − β_B) and fits β by maximum likelihood (logistic regression on battle outcomes with ±1 indicator features). Scores are rescaled to the Elo convention R = 400·β/ln 10 + const, so P = 1/(1 + 10^{(R_B−R_A)/400}). A 50-point gap ≈ 57% win rate, 100 ≈ 64%. Batch BT fitting with bootstrap CIs is preferred over online Elo, which depends on the order of games. Caveats: style confounds (handled with length and markdown covariates), non-uniform sampling, selective disclosure of private variants, and the arena prompt distribution not matching yours.

Log-likelihood vs generative scoring for multiple-choice benchmarks: trade-offs?

Log-likelihood scoring compares log p(option | prompt) across options. It's deterministic, cheap (no decoding) and works for base models without instruction following, which makes it ideal during pretraining. But it needs a length-normalization choice, can't capture reasoning, and isn't available for closed APIs without logprobs. Generative scoring lets the model reason and answer freely, matching real use, but depends on answer extraction, adds sampling variance and costs more. Rankings can differ between the two. Small models often look better under cloze scoring because the "answer with a letter" format is itself an emergent skill.

How does test-set contamination happen, and how would you detect it with and without access to training data?

Benchmarks leak via GitHub, Hugging Face, papers, forums, paraphrases and synthetic data from teachers that saw them, plus adaptive overfitting through checkpoint and mix selection. With training data: n-gram or substring overlap search (for example 13-gram or 50-character matches), embedding similarity for paraphrases. Without it: membership inference such as Min-K% Prob, the exchangeability test (canonical order more likely than shuffled), completion probes, perturbation tests (change numbers or names), and comparing against fresh mirror sets (GSM1k) or post-cutoff items (LiveCodeBench). Prevention: canary strings, private splits, live refresh.

Why did benchmarks like MMLU and SWE-bench Verified stop being useful?

MMLU saturated: frontier models exceed 90%, and with an estimated ~6.5% label-error rate the remaining headroom is mostly noise. Years of public exposure also make contamination likely. SWE-bench Verified: OpenAI reported in Feb 2026 that many tests were overly narrow or checked things the issue never specified, and that frontier models could reproduce gold patches. Scores therefore reflected memorization and test idiosyncrasies more than coding skill, and they recommended SWE-bench Pro. The general lesson is Goodhart plus lifecycle: every public benchmark decays under optimization pressure, so you need successors, private splits and live data.

Design an LLM-as-judge pipeline for comparing two chat models. How do you make it trustworthy?

Use pairwise comparisons on an in-domain prompt set, with a decomposed rubric (correctness, instruction following, safety, concision) and reference answers where possible. Mitigate position bias by judging both orders and counting a win only if it's consistent. Mitigate length bias with length-controlled analysis or rubric penalties. Mitigate self-preference by using a judge from a different family, or a jury of 3 diverse judges. Validate on ~300–500 human-labelled pairs, measure agreement and κ against the human–human ceiling, then freeze the judge version. Report win rates with CIs, analyse by category, and re-validate whenever the judge changes. Never use the same judge as both the RL reward and the final evaluation.

How should you compare two reasoning models fairly?

Compare curves, not points. Sweep reasoning effort or thinking budget and plot accuracy against tokens (or $ or latency), separating sequential (longer CoT) from parallel (maj@k / best-of-N) scaling. Compare at matched budgets and report where each plateaus. Count hidden reasoning tokens in the cost, score truncated outputs as failures, and report the truncation rate. Watch for inverse scaling where more thinking hurts. Use enough samples per problem on small sets like AIME, and use fresh, post-cutoff problems.

What makes a dangerous-capability eval different from a normal capability benchmark?

It's meant to support a claim of the form "the model can't do X", so it must approximate an upper bound: maximal elicitation (best scaffolds, tools, many attempts, fine-tuning away refusals). Under-elicitation produces false negatives, which are the costly error. It measures uplift relative to baselines (internet, experts), sometimes through human uplift trials. Results trigger pre-committed mitigations. Items are often private or hazardous. Recent concerns include sandbagging and evaluation awareness, which can make models look safer in tests than in deployment.

What is a linear probe, and what are the two biggest pitfalls in interpreting one?

A logistic-regression (or similarly simple) classifier trained on frozen activations from a layer to predict some property. High accuracy means the property is linearly decodable there. Pitfall 1: decodable doesn't mean used. The model might never read that direction, so you need causal tests (ablate or steer along the probe direction). Pitfall 2: probe capacity and confounds. An expressive probe can learn the task itself (use control tasks and report selectivity), and the dataset may confound the target with surface features, so test out of distribution. In practice, linear probes are strong, cheap monitors and often beat fancier SAE-based classifiers.

What is the logit lens, why does it work, and when does it fail?

Apply the final layer norm and unembedding to an intermediate residual stream to decode a "current best guess" at each layer. It works because the residual stream is additive and components write in a basis roughly shared with the output. It fails when intermediate representations are in a rotated or shifted basis (often in early layers, and in some model families), and because intermediate layers store information for later computation that isn't meant to be read by the unembedding. The tuned lens fixes much of this with a learned per-layer affine map trained to match the final distribution.

Explain superposition and why it motivates sparse autoencoders.

Networks need to represent far more features than they have neurons. When features are sparse (rarely co-active), the network can assign them nearly orthogonal directions in a lower-dimensional space and tolerate small interference, so a single neuron participates in many features (polysemanticity). Toy models show this happens predictably as sparsity increases. The neuron basis is therefore the wrong unit of analysis. SAEs try to recover the overcomplete feature basis by learning a wide, sparse code that reconstructs activations, in the hope that each latent corresponds to one interpretable feature.

Write down the SAE objective. What is "shrinkage" and how do TopK/JumpReLU address it?

f = ReLU(W_e(x − b_d) + b_e), x̂ = W_d f + b_d, and L = ‖x − x̂‖² + λ Σ_i f_i‖d_i‖ (L1 weighted by decoder norms so the model can't cheat by shrinking f and growing W_d). The L1 penalty pushes all active latents toward zero, so the reconstruction systematically underestimates feature magnitudes. That's shrinkage, and it also blurs the sparsity/fidelity trade-off. TopK SAEs keep only the k largest pre-activations, which sets L0 = k directly with no L1 penalty on magnitudes. JumpReLU learns a per-latent threshold and penalizes L0 using straight-through gradients. Both improve the reconstruction–sparsity frontier.

What are the main limitations of SAEs?

Feature splitting (granularity depends on dictionary size, so there's no canonical feature set), absorption (general features fail to fire when specific ones absorb them), dark matter (a persistent reconstruction error, and splicing in the SAE noticeably increases LM loss), interpretable latents that aren't necessarily causally used, high training cost per layer, and mixed downstream results. GDM found SAE probes worse than linear probes on OOD harmful-intent detection, and synthetic studies find low recovery of ground-truth features. A fair framing: SAEs are good for discovery and hypothesis generation, and weaker than simple baselines for acting on known concepts.

How does activation patching work? Contrast noising and denoising.

Run clean and corrupted inputs that differ in one key factor. Copy an internal activation from one run into the other and measure the change in a metric such as the logit difference of the correct vs. incorrect answer. Denoising (clean into corrupted) tests sufficiency: does restoring this component recover the behaviour? Noising (corrupted into clean) tests necessity: does breaking it destroy the behaviour? Sweeping over layers and positions localizes computation (ROME's causal tracing found mid-layer MLPs at the last subject token for factual recall). Pitfalls: off-distribution corruptions (Gaussian noise), metric choice, and self-repair by backup components.

Describe the induction-head circuit and why it matters.

Two heads in different layers. A previous-token head writes "the previous token was A" into each position's residual. An induction head at a later occurrence of A queries for positions whose previous-token feature equals A (K-composition), attends to the token after the earlier A, and copies it through its OV circuit, predicting B in "…A B … A". Induction heads appear during a sharp phase change in training that coincides with a jump in in-context learning, and Olsson et al. argue they're a major mechanism of ICL. They're the cleanest example of a multi-layer circuit found through weights and causal experiments together.

What are attribution graphs, and how do they differ from SAE feature browsing?

SAEs give you a list of active features. Attribution graphs show how features cause each other on a specific prompt. Anthropic trains a cross-layer transcoder that replaces the MLPs, freezes attention and norms for the prompt, and adds error nodes so the local replacement reproduces the output exactly. The model is then linear in the features, so direct contributions between features can be computed as edges. The graph is pruned and grouped into supernodes, and hypotheses are validated by intervening on features in the real model. It revealed multi-hop reasoning and planning inside a forward pass. Limitations: it explains only a fraction of prompts well, doesn't explain attention formation by default, and depends on transcoder fidelity (error nodes).

How do steering vectors work and what are their failure modes?

Compute a direction associated with a behaviour, typically the mean difference of residual activations between contrastive prompts (CAA), or the top PCA component of the differences (RepE). Add α·v at one or more layers during inference to induce the behaviour, or subtract it to suppress it. You can also project onto v to monitor. They work because many high-level concepts are roughly linear in the residual stream. The refusal direction is a striking case: ablating it removes refusals. Failure modes: too large an α degrades fluency and capability, effects vary across prompts and layers, vectors capture dataset confounds, and steering may fail for concepts that aren't linear. Always measure off-target capability regressions.

Is a reasoning model's chain of thought a faithful explanation? What's the evidence, and does it matter?

Often not. Turpin et al. showed biased few-shot contexts change answers while CoTs rationalize without mentioning the bias. Lanham et al.'s truncation and corruption tests show variable reliance on the CoT. Anthropic's 2025 study found reasoning models verbalize hints they used often below 20% of the time, and RL-learned reward hacks were almost never verbalized. It matters because CoT monitoring is an attractive safety tool. Monitoring can still work when the task requires explicit reasoning, but optimizing CoTs to look good, or moving reasoning into latent space, could make CoTs uninformative. Hence the recommendation to preserve monitorability and not train directly against CoT content.

How do ROME and MEMIT edit facts, and why aren't they widely used for knowledge updates in production?

ROME locates factual recall in a mid-layer MLP via causal tracing, views the MLP's output weights as a key→value memory, and applies a closed-form rank-one update so the subject key maps to a new value with minimal change to other keys. MEMIT distributes updates over several layers for thousands of edits. In production they're rarely used because edits don't propagate to logically related facts (ripple effects), many sequential edits degrade the model, specificity failures are hard to audit, and RAG or fine-tuning give more controllable, reversible knowledge updates.

Design: you're shipping a new frontier model checkpoint next month. What does your eval plan look like?

Layer it. (1) Regression suite: fixed-protocol harness (lm-eval or Inspect) with per-domain loss/BPB plus core capability sets (knowledge, math, code, long context, multilingual), stored per item for paired comparison against the last release. (2) Frontier and agentic: SWE-bench Pro, Terminal-Bench 2.0, τ²-bench with pass^k, OSWorld, HLE/FrontierMath, with scaffolds and budgets documented and accuracy-vs-tokens curves. (3) Fresh or private: post-cutoff items and internal held-out sets to control contamination. (4) Preference: blind pairwise human evals plus a validated judge, style-controlled. (5) Safety: dangerous-capability evals with maximal elicitation, red teaming, jailbreak ASR, propensity evals. (6) Rigor: CIs, paired tests, multiple-comparison control, transcript review of regressions, and a decision rule agreed in advance.