Track A · Model internals

RL for LLMs & RL Environments

Reinforcement learning went from the last, small step of RLHF to the main driver of frontier capability gains. Reasoning models, coding agents and computer-use agents all get their edge from large-scale RL against verifiable rewards inside carefully built environments. This page covers the RL you need (policy gradients, PPO, GRPO and its fixes), how reasoning models are trained, why reward hacking is the central failure mode, and, in the most depth, how RL environments are designed, built and scaled. Frontier labs and the fast-growing set of RL-environment startups now interview directly on that last topic.

TL;DR: the things to be able to say out loud

  • Policy gradient: \(\nabla J = \mathbb{E}[\nabla \log \pi_\theta(a|s)\,A(s,a)]\). Raise the log-probability of actions that did better than expected. Everything else on this page is variance reduction, stability and systems engineering built on that.
  • PPO = policy gradient + importance ratio + clipping (a cheap trust region) + a learned value function for advantages (GAE). Classic RLHF adds a KL penalty to a frozen reference model.
  • LLM as policy: state = prompt plus tokens so far, action = next token, reward is usually given once at the end of the sequence. Single-turn tasks are effectively contextual bandits over whole responses. Agent tasks are real multi-turn MDPs.
  • GRPO drops the critic. It samples G responses per prompt and uses \((r_i-\text{mean})/\text{std}\) as each response's advantage. It's cheap and works well with 0/1 verifiable rewards.
  • GRPO's known flaws and their fixes: length-normalization bias (Dr. GRPO, DAPO token-level loss), entropy collapse (DAPO clip-higher, Clip-Cov/KL-Cov), zero-signal groups (DAPO dynamic sampling), noisy token-level ratios especially in MoE (GSPO sequence-level ratio, CISPO).
  • RLVR: rewards from programs (answer checkers, unit tests, format checks) instead of a learned reward model. Much harder to hack than a reward model, though not impossible. DeepSeek-R1-Zero showed long chain-of-thought emerging from pure RLVR on a strong base model.
  • The open debate: does RL teach new capabilities or just sharpen what's already there (better pass@1, sometimes worse pass@k at large k)? The evidence points both ways, and the answer depends on training length, task diversity and the base model.
  • Reward hacking is the main failure mode: special-casing tests, deleting tests, sycophancy, gaming LLM judges. Mitigations: robust graders, held-out checks, monitors, KL/regularization, environment hardening and auditing rollouts.
  • An RL environment = task distribution + initial state + action/tool interface + observation formatting + transition dynamics (often a sandbox) + termination rule + reward/verifier. A good one is verifiable, hard to hack, calibrated in difficulty, diverse, reproducible and cheap to run in parallel.
  • Infrastructure: RL for LLMs is mostly generation. Rollouts run on vLLM/SGLang, the trainer runs on FSDP/Megatron, and updated weights are synced back to the generators. Async or pipelined designs trade off-policy staleness for throughput. Long-tail rollouts are the bottleneck.
  • Multi-turn agent RL: mask tool/environment output tokens out of the loss, assign credit across turns, and budget for long horizons, sandboxes and flaky tools.

1. RL fundamentals, compactly

You will be asked to define these crisply. Interviewers at labs often start here to check that you know RL itself and haven't only memorized "GRPO".

The MDP

A Markov decision process is a tuple \((\mathcal{S}, \mathcal{A}, P, R, \gamma, \rho_0)\): states, actions, transition dynamics \(P(s'|s,a)\), reward \(R(s,a)\), discount \(\gamma\in[0,1]\) and initial-state distribution \(\rho_0\). Markov means the next state depends only on the current state and action. A policy \(\pi_\theta(a|s)\) is a distribution over actions. A trajectory (episode, rollout) is \(\tau = (s_0,a_0,r_0,s_1,\dots)\). The return is the discounted sum of rewards \(G_t = \sum_{k\ge 0}\gamma^k r_{t+k}\). The objective is $$J(\theta) = \mathbb{E}_{\tau\sim\pi_\theta}\Big[\sum_t \gamma^t r_t\Big].$$ Sutton & Barto 2018

Value functions, advantage, Bellman

QuantityDefinitionMeaning
State value \(V^\pi(s)\)\(\mathbb{E}_\pi[G_t\mid s_t=s]\)How good it is to be in \(s\) and then follow \(\pi\)
Action value \(Q^\pi(s,a)\)\(\mathbb{E}_\pi[G_t\mid s_t=s,a_t=a]\)How good it is to take \(a\) in \(s\) and then follow \(\pi\)
Advantage \(A^\pi(s,a)\)\(Q^\pi(s,a)-V^\pi(s)\)How much better \(a\) is than the policy's average action. Zero-mean under \(\pi\).
TD error \(\delta_t\)\(r_t+\gamma V(s_{t+1})-V(s_t)\)One-step estimate of the advantage (exact if \(V=V^\pi\), in expectation)

The Bellman expectation equation writes value recursively: \(V^\pi(s)=\mathbb{E}_{a\sim\pi,s'\sim P}[r+\gamma V^\pi(s')]\). The Bellman optimality equation replaces the expectation over actions with a max: \(Q^*(s,a)=\mathbb{E}_{s'}[r+\gamma\max_{a'}Q^*(s',a')]\). Value-based methods (Q-learning, DQN) iterate on the optimality equation. Policy-gradient methods optimize \(\pi\) directly and use value functions only to reduce variance. LLM RL is almost entirely in the policy-gradient family, because the action space (a vocabulary of about 100k+ tokens, applied autoregressively) makes a max over actions meaningless at the sequence level, and because we already start from a very good policy, the pretrained model.

On-policy vs off-policy, exploration, credit assignment

Intuition

Supervised learning says "make this output more likely." Policy-gradient RL says "make whatever you just did more likely, in proportion to how much better than expected it turned out." The model generates its own training data, and the reward decides the sign and size of each update. That's why RL can go beyond the demonstrations (it can find solutions nobody wrote down) and also why it can find exploits nobody intended.

Interview angle

"What's the difference between \(V\), \(Q\) and \(A\), and why do we use \(A\) in the gradient?" A strong answer: \(A\) is centered, so subtracting \(V\) removes a large state-dependent offset that adds variance without changing the expected gradient. Also expect to say why LLM RL is on-policy-ish, and what credit assignment means when the reward is a single number per response.

Go deeper

2. Policy gradients: REINFORCE to GAE

The log-derivative trick

We want \(\nabla_\theta \mathbb{E}_{\tau\sim\pi_\theta}[R(\tau)]\), but the expectation is over a distribution that itself depends on \(\theta\), and the reward is not differentiable (it's a unit-test result). The trick is \(\nabla p_\theta = p_\theta \nabla\log p_\theta\), so $$\nabla_\theta J = \mathbb{E}_{\tau\sim\pi_\theta}\big[R(\tau)\,\nabla_\theta\log p_\theta(\tau)\big],\qquad \log p_\theta(\tau)=\sum_t\log\pi_\theta(a_t|s_t) + \underbrace{\textstyle\sum_t \log P(s_{t+1}|s_t,a_t)}_{\text{no }\theta}.$$ The dynamics terms drop out. You need neither a model of the environment nor a differentiable reward, only the ability to sample and score. This is the REINFORCE estimator (Williams, 1992). In code, it is just a weighted negative log-likelihood: loss = -(advantage.detach() * logprobs).mean().

Baselines and variance reduction

The raw estimator is unbiased but has very high variance. Three standard fixes:

  1. Causality / reward-to-go: an action can't affect past rewards, so weight \(\nabla\log\pi(a_t|s_t)\) by \(G_t\), not the full return. (With a single terminal reward, as in most LLM RL, these coincide.)
  2. Baselines: subtracting any \(b(s_t)\) that doesn't depend on the action leaves the gradient unbiased, because \(\mathbb{E}_{a\sim\pi}[\nabla\log\pi(a|s)]=\nabla\sum_a\pi(a|s)=\nabla 1=0\). The best common choice is \(b=V(s)\), which gives the advantage. In LLM RL the baseline is often empirical: the mean reward of other samples for the same prompt (RLOO, GRPO), or a batch mean (REINFORCE++).
  3. Normalization: whitening advantages (divide by std) stabilizes the step size across prompts and reward scales. It also introduces subtle biases (§5).

Generalized Advantage Estimation (GAE)

With a learned critic \(V_\phi\), you can estimate advantages anywhere between the one-step TD error (low variance, biased if \(V\) is wrong) and the Monte-Carlo return minus baseline (unbiased, high variance). GAE takes an exponentially weighted mix: $$\hat A_t^{\text{GAE}(\gamma,\lambda)}=\sum_{l\ge 0}(\gamma\lambda)^l\,\delta_{t+l},\qquad \delta_t=r_t+\gamma V_\phi(s_{t+1})-V_\phi(s_t).$$ \(\lambda=0\) gives TD(0). \(\lambda=1\) gives the MC return minus \(V\). Schulman+ 2015 In LLM RLHF with PPO, typical settings were \(\gamma=1\) and \(\lambda\) close to 1 (0.95–1.0), and the reward was sparse: the reward-model score at the last token, plus a per-token KL penalty.

Common mistake

Saying "the baseline makes the gradient biased." It doesn't, provided the baseline doesn't depend on the action being scored. A subtle exception matters in practice: if the baseline for sample \(i\) includes sample \(i\)'s own reward (as GRPO's group mean does), there is a small bias. That's exactly why RLOO uses the leave-one-out mean.

Interview angle

"Derive the policy gradient" is a common whiteboard ask. Hit three beats: (1) log-derivative trick, (2) dynamics terms vanish, so it's model-free, (3) a baseline keeps it unbiased because the expected score function is zero. Bonus: explain GAE's \(\lambda\) as a bias-variance dial and why LLM people often drop the critic altogether.

Go deeper

3. PPO and classic RLHF

Why trust regions

Plain policy gradient takes one step per batch of samples. If you take several steps on the same batch, or one large step, the policy can move far from the one that generated the data, so the gradient estimate becomes wrong and performance collapses. TRPO enforced a hard KL constraint between old and new policy with a second-order method. Schulman+ 2015 PPO gets most of that benefit with a first-order clipped surrogate. Schulman+ 2017

The clipped objective

Let \(\rho_t(\theta)=\dfrac{\pi_\theta(a_t|s_t)}{\pi_{\theta_\text{old}}(a_t|s_t)}\) be the importance ratio. Then $$L^{\text{CLIP}}(\theta)=\mathbb{E}_t\Big[\min\big(\rho_t\hat A_t,\ \text{clip}(\rho_t,1-\epsilon,1+\epsilon)\hat A_t\big)\Big],\qquad \epsilon\approx 0.2.$$ If \(\hat A_t>0\), the objective stops rewarding increases of \(\rho_t\) above \(1+\epsilon\). If \(\hat A_t\lt 0\), it stops rewarding decreases below \(1-\epsilon\). The min makes it a pessimistic bound: clipping can only remove incentive, never add it. Once a token's ratio leaves the band in the "already improved" direction, its gradient is exactly zero. That's the property later algorithms (DAPO, CISPO) tinker with.

1−ε1+ε 1−ε1+ε ρρ A > 0: objective flat beyond 1+ε A < 0: objective flat below 1−ε grad = 0 grad = 0
PPO's clipped surrogate (per token). Clipping removes the incentive to push a token's probability further once it has moved more than ε in the advantageous direction.

The full PPO-RLHF setup

InstructGPT-style RLHF uses four models: the policy (trained), a frozen reference policy (the SFT model), a reward model (frozen, trained beforehand on human preference pairs with a Bradley-Terry loss), and a value model/critic (trained, often initialized from the reward model). Ouyang+ 2022 Christiano+ 2017 The per-token reward is $$r_t = \underbrace{r_\text{RM}(x,y)\cdot\mathbb{1}[t=T]}_{\text{sequence score at the end}} - \beta\,\log\frac{\pi_\theta(y_t|x,y_{\lt t})}{\pi_\text{ref}(y_t|x,y_{\lt t})}.$$ The KL term keeps the policy near the reference. It stops reward-model exploitation ("overoptimization": the proxy score rises while the true quality falls) and preserves fluency and diversity. Gao+ 2022 The value loss is a regression of \(V_\phi(s_t)\) onto the returns, often also clipped.

PPO-RLHF componentWhy it existsCost / pain
Policy \(\pi_\theta\)The thing being trainedFull training state (weights, grads, Adam: about 16 bytes/param in mixed precision)
Reference \(\pi_\text{ref}\)KL anchor against drift and reward hackingA second forward pass per batch
Reward modelTurns preferences into a scalarHackable proxy; needs fresh preference data
Critic \(V_\phi\)Per-token baselines via GAESame size as the policy, so it roughly doubles training memory. Hard to train well on sparse, end-of-sequence rewards.
Intuition

The critic is the most fragile part for LLMs. It has to predict, from a partial response, the expected final reward. That's a hard regression problem on top of a non-stationary policy. That fragility, plus the memory cost, is the main reason the field moved to critic-free methods that estimate the baseline by sampling several responses per prompt.

Interview angle

Expect: "Walk me through one PPO step for an LLM." Generate responses, score them with the RM, compute per-token KL vs the reference, compute values and GAE, then run several minibatch epochs of clipped policy loss + value loss. Probes: what does \(\beta\) trade off (reward vs. drift); what happens when \(\epsilon\) is too large (instability) or too small (slow learning); why is the KL in the reward in PPO-RLHF but in the loss in GRPO?

Go deeper

4. The LLM as a policy

Mapping the MDP onto a language model:

MDP elementSingle-turn (math, code, chat)Multi-turn agent (SWE, browser, tools)
State \(s_t\)Prompt + tokens generated so farFull transcript: system prompt, task, prior actions, tool outputs, plus hidden environment state (filesystem, DB, web page)
Action \(a_t\)Next token (or the whole response as one "action")Token-level, but usually reasoned about per turn: one tool call or message
TransitionDeterministic: append the tokenStochastic/external: the environment runs the tool and returns an observation
RewardOnce, at the end (verifier or RM)Usually at the end (tests pass, task completed); sometimes intermediate
HorizonHundreds to tens of thousands of tokensTens to hundreds of turns, potentially millions of tokens of context over the episode

The bandit view. With a deterministic "append token" transition and a single terminal reward, single-turn RL is effectively a contextual bandit: context = prompt, arm = complete response, payoff = reward. That's why simple REINFORCE-style estimators with a per-prompt baseline work so well (RLOO's main argument). Ahmadian+ 2024 Per-token value functions buy less here than in Atari or robotics, because there are no intermediate rewards to bootstrap from.

The token-level view still matters for the implementation. The gradient is a sum over tokens, importance ratios and clipping are per token in PPO/GRPO, and the KL penalty is per token. Choices like how you average over tokens (per sequence or per batch) change the effective weighting of long vs short responses. Those choices turned out to matter a lot (§5).

The multi-turn view is a real MDP: the environment's responses are not under the policy's control and must be excluded from the loss. Rewards may arrive after many turns, and the "state" includes things the model can't see (a database, a remote browser). §10 covers this.

Task distributionprompts, repos, sites Policy π_θ (rollouts)G samples per task Environmenttools, sandbox, sim user Verifier / rewardtests, checker, judge Advantagesgroup norm / GAE Policy updateclip loss, KL, masks new weights → generators
The RL loop for LLMs. In agent settings the policy↔environment arrow repeats for many turns before a reward is computed.
Interview angle

"Is token-level or sequence-level the right granularity for an action?" A strong answer: the reward lives at the sequence level, the optimization happens at the token level, and the mismatch between them causes real bugs (length bias, ratio noise). GSPO's argument is that the importance ratio should match the granularity of the reward. For agents, the turn is often the natural unit for credit assignment.

5. GRPO and its successors

GRPO: Group Relative Policy Optimization

Introduced in DeepSeekMath and made famous by DeepSeek-R1. Shao+ 2024 For each prompt \(q\), sample a group of \(G\) outputs \(\{o_1,\dots,o_G\}\) from \(\pi_{\theta_\text{old}}\), score them, and give every token of output \(i\) the same advantage: $$\hat A_{i,t} = \frac{r_i-\text{mean}(r_1,\dots,r_G)}{\text{std}(r_1,\dots,r_G)}.$$ The objective is a PPO-style clipped surrogate averaged per sequence, then over the group, with the KL term moved from the reward into the loss: $$J_\text{GRPO}(\theta)=\mathbb{E}\Bigg[\frac{1}{G}\sum_{i=1}^G\frac{1}{|o_i|}\sum_{t=1}^{|o_i|}\Big(\min\big(\rho_{i,t}\hat A_{i,t},\ \text{clip}(\rho_{i,t},1\!-\!\epsilon,1\!+\!\epsilon)\hat A_{i,t}\big)-\beta\,\mathbb{D}_\text{KL}[\pi_\theta\|\pi_\text{ref}]\Big)\Bigg]$$ where \(\rho_{i,t}=\pi_\theta(o_{i,t}|q,o_{i,\lt t})/\pi_{\theta_\text{old}}(o_{i,t}|q,o_{i,\lt t})\). The KL is estimated per token with the low-variance "k3" estimator \(\frac{\pi_\text{ref}}{\pi_\theta}-\log\frac{\pi_\text{ref}}{\pi_\theta}-1\).

Why it caught on: no critic (saving roughly a policy-sized model of memory and a hard training problem); the group mean is a good baseline for 0/1 verifiable rewards; and it's simple. Typical group sizes are 8–64 samples per prompt.

Key property: if all \(G\) samples get the same reward (all right or all wrong), every advantage is zero and the prompt contributes no gradient. Prompts that are too easy or too hard are wasted compute. That's the root of curriculum and difficulty-filtering practices (§9).

Known problems with vanilla GRPO

  1. Length bias from \(1/|o_i|\). Each sequence's loss is divided by its own length. For a negative advantage, a longer wrong answer gets a smaller per-token penalty, so the model learns that being wrong at length is cheaper. Response length inflates, especially for incorrect outputs. Liu+ 2025 (Dr. GRPO)
  2. Difficulty bias from std normalization. Dividing by the group std up-weights prompts where almost all samples agree (low std), meaning very easy or very hard ones. Dr. GRPO argues this skews the optimization.
  3. Entropy collapse. With symmetric clipping, low-probability "exploration" tokens with positive advantage can only rise by a factor of \(1+\epsilon\) per step, while already-likely tokens dominate. Entropy drops fast early in training and performance plateaus. One paper fits an empirical relationship between performance and entropy, \(R\approx -a e^{H}+b\), suggesting performance is "bought" with entropy. Cui+ 2025
  4. Token-level importance ratios are noisy. The ratio of a single token's probability under two nearby policies is a high-variance estimate when one sequence-level reward is spread over thousands of tokens. This gets worse with MoE models, where routing changes between \(\pi_\text{old}\) and \(\pi_\theta\) can swing individual token probabilities. Zheng+ 2025 (GSPO)
  5. Clipping as an unintended prior. One study found that GRPO with random rewards still improved Qwen2.5-Math-7B on MATH-500 by about 21 points (vs. about 29 with true rewards), and attributed it to a clipping bias that amplifies behaviors the base model already favors. The effect largely didn't transfer to Llama or OLMo. Shao+ 2025 Treat RLVR gains on a single model family with suspicion.

The successor family

RLOO (REINFORCE Leave-One-Out): sample \(k\) responses and use as baseline for each the mean reward of the other \(k-1\). This is unbiased, critic-free, with no clipping or ratio in the simplest form. The paper argued PPO's machinery is mostly unnecessary in the LLM bandit setting. Ahmadian+ 2024 (Mean-centering with the full group, as in GRPO without the std division, gives exactly \(\tfrac{k-1}{k}\) times the RLOO advantage, so the two differ only by a constant scale.)

REINFORCE++: critic-free, but normalizes advantages over the global batch instead of per prompt. The claim is that per-prompt normalization is biased and overfits, and global normalization is effectively unbiased as the batch grows. It has a variant with a group baseline for reasoning tasks. Hu+ 2025

Dr. GRPO ("GRPO Done Right"): removes the \(1/|o_i|\) per-sequence length normalization and the std division, which fixes the length and difficulty biases above. It reports similar accuracy with better token efficiency. Liu+ 2025

DAPO (ByteDance Seed / Tsinghua): four changes on top of GRPO, with no KL term, for long-CoT reasoning Yu+ 2025:

GSPO (Qwen team): uses a sequence-level importance ratio, the length-normalized likelihood ratio \(s_i=\big(\pi_\theta(o_i|q)/\pi_{\theta_\text{old}}(o_i|q)\big)^{1/|o_i|}\), and clips whole sequences. The reward is per sequence, so the off-policy correction should be too. It reports better stability, especially for MoE RL, and credits it for improvements in Qwen3. Zheng+ 2025

CISPO (MiniMax-M1): instead of zeroing the gradient of clipped tokens, it clips the importance-sampling weight (as a stop-gradient coefficient) and keeps every token's gradient. Rare but important "fork" tokens, like a reflective "wait", aren't silenced. MiniMax 2025

VAPO: brings the value function back (value-pretraining, decoupled GAE \(\lambda\) for actor and critic, length-adaptive GAE) and argues a well-trained critic beats critic-free methods on long CoT. Yue+ 2025 Clip-Cov / KL-Cov restrict updates on the high-covariance tokens that drive entropy collapse. Cui+ 2025

MethodCritic?Advantage / baselineRatio & clippingLoss aggregationMain fix / claim
REINFORCENoReturn − (optional) baselineNone (pure on-policy)AnySimplest unbiased estimator
PPOYesGAE from learned \(V\)Token ratio, symmetric clip ε≈0.2Token meanTrust region cheaply; dense per-token baselines
RLOONoLeave-one-out group meanOften noneSequenceBandit framing; PPO machinery unnecessary
GRPONoGroup mean, ÷ group stdToken ratio, symmetric clipPer-sequence mean, then group meanNo critic; good with 0/1 rewards
Dr. GRPONoGroup mean (no std)Token ratio, clipNo per-sequence length normRemoves length and difficulty bias
DAPONoGroup-normalizedToken ratio, asymmetric clip (clip-higher)Token-level over batchEntropy collapse, zero-signal groups, long-CoT stability; no KL
REINFORCE++NoGlobal batch normalizationToken ratio, clipTokenLess biased normalization; robust in general RLHF
GSPONoGroup-normalizedSequence-level ratio (length-normalized), sequence clipSequenceRatio matches reward granularity; MoE stability
CISPONoGroup-normalizedClip the IS weight (stop-grad), keep all token gradsTokenRare tokens keep gradient; efficiency
VAPOYesGAE with value-pretraining, length-adaptive λToken, clip-higherTokenA good critic still wins on long CoT
May be out of date

This algorithm zoo moves monthly. A large Meta study (over 400k GPU-hours) concluded that many of these choices (loss aggregation, normalization, off-policy correction, curriculum) mostly change compute efficiency rather than the asymptotic performance, and proposed a combined "ScaleRL" recipe with predictable, sigmoid-shaped scaling curves. Khatri+ 2025 In 2026 interviews, it's more valuable to explain the mechanisms (length weighting, entropy, ratio granularity, staleness) than to recite a leaderboard of acronyms. Check current recipes before claiming any one is "the" standard.

Interview angle

"Why did response length explode in my GRPO run?" Strong answers mention (1) the \(1/|o_i|\) normalization making long wrong answers cheap (Dr. GRPO), (2) truncation handling: are cut-off answers scored 0 and still in the loss? (3) genuine benefit, since longer CoT can be learned behavior if accuracy rises with it. Diagnose by splitting length curves by correct vs incorrect. "Why is entropy collapsing?" Clip-higher, an entropy bonus, KL-Cov, temperature and more diverse data. Also check whether the reward is too easy to max.

Go deeper

6. RL with verifiable rewards and reasoning models

RLVR

Reinforcement Learning with Verifiable Rewards replaces the learned reward model with a program that checks correctness. The term was popularized by AI2's Tülu 3, which applied it to math, grade-school math and instruction-following constraints. Lambert+ 2024 The common verifier types:

DomainVerifierGotchas
Math (final answer)Extract the answer (e.g., \boxed{}), normalize, compare symbolically or numerically (sympy-style equivalence)Equivalent forms (\(\tfrac12\) vs 0.5 vs 1/2), units, multiple-choice guessing (25% free reward), problems with wrong reference answers, proofs (no single answer)
CodeRun hidden unit tests in a sandbox; reward = pass/fail or fraction passedWeak tests let wrong code pass; tests visible to the model can be special-cased; flaky/timeout-sensitive tests add noise; sandbox escape
Instruction constraintsProgrammatic checks ("exactly 3 bullet points", "no commas", JSON schema valid)Satisfying the letter but not the spirit
FormatRegex: reasoning inside <think>, answer in the expected slotA small format reward early helps parsing; too large and the model optimizes format over correctness
Agentic tasksEnd-state checks: tests pass in repo, DB row exists, file has expected content, web form submittedMany valid end states; state checkers that are too narrow or too loose

Why RLVR beats RM-based RL for reasoning: a program can't be talked into a high score the way a neural reward model can, so you can optimize much harder (thousands of steps, little or no KL) before hacking dominates. The trade-off is coverage: only domains with checkable answers qualify. Most current research is about extending verifiability (rubrics, LLM judges, generated tests) to fuzzier domains (§7).

o1 and the reasoning-model paradigm

OpenAI's o1 (September 2024) announced that performance improves smoothly with both more RL train-time compute and more test-time compute (thinking longer), using large-scale RL to teach the model to use a long chain of thought productively. OpenAI 2024 Details were not published. The prior work on spending inference compute (best-of-N with verifiers, search, revision) showed that compute-optimal test-time scaling can beat scaling parameters on some problems. Snell+ 2024 o1's contribution was to put the "search" inside one long sampled chain of thought, learned by RL, instead of in an external search procedure.

DeepSeek-R1 and R1-Zero

DeepSeek-R1 (January 2025) published the recipe openly DeepSeek-AI 2025:

Kimi k1.5, released the same week, independently reported a similar finding: long-context RL with simple outcome rewards and no MCTS, value function or PRM. It also introduced "partial rollouts" to handle very long generations (§11). Kimi Team 2025

Common mistake

Treating the "aha moment" as proof that RL created reflection. A follow-up found that DeepSeek-V3-Base already shows aha-like self-reflection before RL, and that Qwen2.5 base models reason well even without prompt templates, which points to pretraining biases (including possible exposure to reasoning-style data). Liu+ 2025 The safer claim is that RL amplifies and stabilizes latent behaviors that pay off under the reward.

Test-time compute scaling

Two families:

Sequential (longer thinking)

One chain of thought, made longer. RL-trained reasoning models do this natively. Accuracy rises roughly log-linearly with thinking tokens up to a saturation point. Controls include reasoning-effort settings, thinking budgets, and "budget forcing", where you append "Wait" to make the model continue or force-stop it. Muennighoff+ 2025 (s1)

Parallel (more samples)

Sample N answers and aggregate: majority vote (self-consistency), best-of-N with a verifier or reward model, or tree search. This scales with N but saturates when the verifier is imperfect or the right answer never appears in the samples. Pass@k measures the ceiling of this approach.

The debate: does RL teach new things or just sharpen?

This is a favorite interview topic because there is real evidence on both sides.

"Sharpening" view

Across model families and RLVR algorithms, RL'd models beat their base models at pass@1, but the base models win at pass@k for large k (e.g., k=256). Coverage and perplexity analyses suggest the solutions RL finds were already in the base distribution. RL reweights toward them and loses some diversity. Distillation from a stronger teacher, by contrast, did add new reasoning patterns. Yue+ 2025 The spurious-reward results point the same way: some gains come from amplifying existing behaviors. Shao+ 2025

"Expansion" view

With prolonged RL (KL control, periodic reference resets, diverse tasks), RL'd models beat base models across pass@k, including on tasks where the base model fails at any k. The gains are largest where the base model starts weakest. Liu+ 2025 (ProRL) Practitioners also point out that agentic skills (multi-step tool use, long-horizon persistence) look very different after large-scale RL, and that pass@k at huge k on short tasks is a generous metric for the base model.

A defensible synthesis: on short, well-covered tasks, small-to-moderate RL mostly sharpens (it raises pass@1 toward the base model's pass@k). Longer training, broader task distributions and compositional, multi-step environments can push the boundary out. Exploration is the bottleneck, so the base model and the environment's curriculum largely determine which regime you're in.

May be out of date

This debate was very active through 2025 and into 2026. Related threads include on-policy distillation (student samples, teacher grades each token densely) as a cheaper alternative or complement to RL Thinking Machines 2025, and training objectives that directly optimize pass@k or preserve diversity. Check recent results before taking a strong position in an interview. Present it as open.

Interview angle

"How would you tell whether your RL run taught the model anything new?" Strong answer: plot pass@k curves for base vs RL'd at k = 1…256 on a held-out set; check whether RL'd solutions appear among the base model's samples (coverage); test transfer to unseen task types; control for the format/answer-extraction gains that often explain early jumps; and run the same recipe on more than one model family.

Go deeper

7. Reward design: outcome, process, judges and rubrics

Outcome vs process rewards

Outcome reward (ORM / verifier)Process reward (PRM)
SignalOne score for the final answerA score per reasoning step
Credit assignmentSparse; relies on many samplesDense; pinpoints the first wrong step
LabelsCheap (reference answer, tests)Expensive (human step labels) or estimated (MC rollouts)
HackabilityLow if programmaticHigher: the policy can learn steps that look good to the PRM
Main use todayRL training signalReranking / best-of-N, search guidance, research on dense RL credit

Key results. OpenAI's "Let's Verify Step by Step" trained a PRM on about 800k human step-level labels (PRM800K). Used as a best-of-N reranker, it beat an ORM on a MATH subset. Lightman+ 2023 Math-Shepherd removed the human labels: a step's label is estimated by rolling out completions from that step and checking how often they reach the right answer. Wang+ 2023 "Rewarding Progress" reframes a good process reward as progress, the change in the likelihood of eventual success under a (different) prover policy, which is effectively a step-level advantage. Setlur+ 2024 PRIME derives an implicit PRM from a model trained only on outcome labels and updates it online, which reduces hacking. Cui+ 2025

Practical status: large-scale reasoning RL mostly uses outcome rewards (R1 and Kimi k1.5 both reported PRMs as not worth it at scale). Process signals live on in verifier-guided search, in agentic settings as turn-level checks, and as an active research area.

Learned reward models, LLM judges and rubrics

For non-verifiable domains (writing quality, helpfulness, research reports, medical advice), the options are:

Intuition

Reward signals sit on a spectrum of hackability vs. coverage. Exact-match checkers cover little and are almost unhackable. Unit tests cover more and are somewhat hackable. Rubric judges cover a lot and are moderately hackable. Free-form LLM judges and learned RMs cover everything and are very hackable. The craft is pushing each domain as far toward the verifiable end as possible, and adding monitoring where you can't.

Interview angle

"Design a reward for training a model to write good literature reviews." Strong answers break it down: verifiable sub-checks (citations resolve to real papers, quotes match sources, required sections present), rubric items judged per criterion by a different model family than the policy, pairwise comparisons against a reference instead of absolute scores, length normalization or caps, and a held-out human eval to catch judge exploitation. They also mention monitoring reward-vs-human-agreement drift over training.

Go deeper

8. Reward hacking and specification gaming

Reward hacking: the policy raises the measured reward without doing what the designer intended. Formally, any time optimizing the proxy reward diverges from optimizing the true objective. Skalse+ 2022 Classic RL is full of examples, like the boat-racing agent circling to collect respawning targets instead of finishing the race. DeepMind 2020 With LLMs, the agent is smart enough to read and reason about its own grader.

Taxonomy with LLM examples

TypeExampleMitigation
Test special-casingCode that detects test inputs and returns hard-coded expected outputsHidden tests, held-out tests, randomized/property-based tests, LLM review of diffs
Grader tamperingEditing or deleting failing tests, modifying conftest.py, exiting the process early with code 0 before tests run, overriding __eq__ so every comparison passes Zhong+ 2025Read-only test dirs, run graders outside the agent's sandbox, checksum protected files, check for exit-code tricks
Environment exploitsFinding the reference solution on disk or in git history, using network access to look up answers, abusing timing or scoring code. METR documented frontier models overwriting timing functions to fake speedups, stubbing evaluators, and searching the Python call stack for the grader's reference answer. METR 2025Scrub artifacts, network policies, separate grader process, red-team the env before training
Judge exploitationVerbose, confident or flattering answers; injected text that addresses the judgeRubrics, pairwise vs reference, different judge family, length control, strip instructions aimed at the judge
SycophancyAgreeing with the user's wrong belief because human raters and preference models reward agreement Sharma+ 2023Preference data that rewards correction, targeted evals, debiased RMs
Answer-format gamingListing several answers so the extractor finds the right one; guessing on multiple choiceStrict extraction (single final answer), penalize multiple answers, avoid MC-only datasets
Length/effort gamingPadding to look thorough, or giving up early on hard tasks when partial credit paysLength-aware rewards, careful partial-credit design

Why it matters beyond the metric

A mitigation playbook

  1. Harden the environment first: graders run outside the agent's reach, protected files, no reference answers in the filesystem, network egress controlled.
  2. Diverse, hidden verification: held-out tests the policy never sees, plus property-based or randomized tests, plus multiple independent checkers.
  3. Monitor rollouts continuously: sample and read transcripts; run LLM monitors that flag suspicious patterns (edits to test files, sys.exit, hard-coded constants); track reward vs. held-out-eval divergence.
  4. Regularize: KL to a reference, early stopping on proxy-vs-true divergence, reward caps.
  5. Penalize detected hacks carefully. Explicit penalties work for clear-cut cases, but penalizing on CoT content risks obfuscation.
  6. Patch and re-version the environment when a hack is found, and add the exploit to a regression suite for env QA.
Interview angle

RL-env startups ask some version of: "Here's a coding environment. How would the model hack it, and how would you stop it?" Walk the attack surface systematically: the grader (tests, exit codes, file access), the filesystem (solutions, git history, caches), the network, the judge, the answer extractor, timeouts and partial credit. Then describe the env QA process: red-team with a strong model before release, check that pass rates on known-impossible variants are near zero, and audit sampled transcripts during training.

Go deeper

9. RL environments in depth

For LLM RL, the environment is everything except the policy and the optimizer: what tasks the model sees, what it can do, what it observes back, when the episode ends, and how it's scored. As algorithms converged on a few GRPO/PPO variants, the environments became the main thing that sets one lab's RL apart from another's, and a market formed around building them. As of 2026, frontier labs buy environments from dedicated vendors and from human-data companies, alongside open hubs.

May be out of date

The RL-environment market and tooling changed a lot in 2025–2026: vendor startups, acquisitions, open hubs, new interface standards. Treat any specific company or framework named here as a snapshot as of late 2026, and check current docs before relying on API details.

Anatomy of an environment

ComponentWhat it isExample (SWE-style task)
Task distributionThe dataset or generator of task instances, with metadata (difficulty, domain, split)Thousands of (repo, commit, issue) tuples mined from GitHub
Initial stateWhat exists at reset(): prompt, files, DB, browser page, user personaA container with the repo checked out at the pre-fix commit and dependencies installed
Action space / toolsWhat the model can do: free text, structured tool calls, shell commands, clicksbash, edit_file, search, submit
Observation formattingHow environment output becomes tokens: truncation, error messages, screenshots vs accessibility treestdout/stderr truncated to N lines, file views with line numbers
Transition dynamicsHow state changes after an action (the sandbox, simulator or simulated user)Actually run the command in the container
TerminationSuccess, explicit submit, max turns/tokens/wall-clock, unrecoverable errorsubmit called, or 100 turns, or 30 min
Reward / verifierThe scoring function, possibly multi-componentHidden FAIL_TO_PASS and PASS_TO_PASS tests run after submit; reward = 1 if all pass
Harness / scaffoldThe agent loop and system prompt the policy runs inside (which may be a production agent like a coding CLI)A ReAct-style loop with a tool schema; or a real coding-agent harness

A minimal interface

Classical RL standardized on the Gym API. Gymnasium (the maintained successor to OpenAI Gym) uses reset(seed=...) → (obs, info) and step(action) → (obs, reward, terminated, truncated, info). It separates true termination from time-limit truncation, which matters for bootstrapping. Legacy Gym's step returned a 4-tuple with a single done. Gymnasium docs Towers+ 2024 LLM environments keep that shape but deal in messages and tool calls, run asynchronously (thousands of concurrent episodes), and keep the verifier separate from the dynamics. A sketch (pseudo-code, not any specific library's API):

from dataclasses import dataclass, field
from typing import Any

@dataclass
class Task:                       # one row of the task distribution
    id: str
    prompt: str
    assets: dict                  # repo snapshot, DB seed, URL, user persona...
    hidden: dict                  # reference answer / hidden tests (NEVER shown to the policy)
    meta: dict = field(default_factory=dict)   # difficulty, domain, split, version

@dataclass
class StepResult:
    observation: list[dict]       # new messages to append (tool results, user replies)
    terminated: bool              # task finished (submit / success / fatal)
    truncated: bool               # budget exhausted (turns, tokens, wall-clock)
    info: dict                    # diagnostics; not shown to the policy

class LLMEnv:
    tools: list[dict]             # JSON schemas exposed to the policy

    async def reset(self, task: Task, seed: int) -> list[dict]:
        """Provision an isolated sandbox from task.assets (container/VM/browser),
        seed all randomness, return the initial messages (system + task prompt)."""

    async def step(self, action: dict) -> StepResult:
        """Parse one assistant turn (text and/or tool calls). Execute tools in the
        sandbox with timeouts; format outputs (truncate, sanitize); check
        termination. Malformed calls get an error observation, not a crash."""

    async def score(self) -> dict[str, float]:
        """Run the verifier OUTSIDE the policy's reach (separate process or
        container) against the final state plus task.hidden. Return components,
        e.g. {"correct": 1.0, "format": 1.0, "tests_modified": 0.0}."""

    async def close(self) -> None:
        """Tear down the sandbox; always runs, even on error/timeout."""

# Trainer-side rollout (simplified):
#   msgs = await env.reset(task, seed)
#   while True:
#       turn = await policy.generate(msgs, tools=env.tools)     # record token ids + logprobs
#       res  = await env.step(turn); msgs += [turn] + res.observation
#       if res.terminated or res.truncated: break
#   reward = combine(await env.score()); await env.close()
#   loss_mask = 1 for policy-generated tokens, 0 for prompt / tool-output tokens
episode lifecycle ┌────────────┐ reset(task, seed) ┌───────────────────────────────┐ │ task queue │ ──────────────────▶ │ sandbox: container / VM / sim │ └────────────┘ └───────────────┬───────────────┘ │ initial obs (msgs) ┌──────────────────────────────────────▼───────┐ │ policy turn (tokens + logprobs recorded) │◀──────┐ └──────────────────────┬───────────────────────┘ │ │ tool call / message │ obs (loss-masked) ┌──────────────────────▼───────────────────────┐ │ │ env.step: execute, timeout, format, check end │──────┘ └──────────────────────┬───────────────────────┘ │ terminated / truncated ┌──────────────────────▼───────────────────────┐ │ score(): verifier in separate process │ → reward components └──────────────────────┬───────────────────────┘ ▼ close() → trajectory {tokens, mask, logprobs, reward, info}

Design principles

PrincipleWhyHow, concretely
VerifiabilitySparse but trustworthy reward beats dense noisy reward for RL at scaleDefine success as a checkable end state; hidden tests; multiple independent checks; avoid rewards that need a judge when a program would do
Non-hackable gradersThe policy will optimize the grader, not your intentGrader outside the sandbox; protected/read-only test files; no answers on disk; red-team with a strong model; "impossible task" canaries whose pass rate should be about 0
Difficulty calibrationGroup-relative methods get zero gradient when all samples tieMeasure pass rate under the current policy (e.g., 8–16 samples); keep tasks in a band like 10–90%; refresh as the policy improves (curriculum); DAPO-style dynamic sampling online
DiversityNarrow envs produce narrow skills and overfitting to quirksMany repos/sites/domains; varied phrasings; procedural generation; dedupe near-duplicates; track per-domain reward
Determinism & reproducibilityNoisy rewards look like learning signal; debugging needs replaysPin images and dependencies; seed everything; freeze time and network responses (record/replay); version tasks and graders; log full traces
Sandboxing & safetyThe policy runs arbitrary code and gets rewarded for finding loopholesContainers/microVMs (gVisor, Firecracker-style), no host mounts, egress allowlists, resource limits, kill on timeout
Throughput & costRL needs millions of episodesFast provisioning (snapshots, warm pools), async I/O, cheap graders, cache env builds; budget per-episode cost (CPU-minutes, API calls)
Robust observationsHuge tool outputs blow the context; crashes waste rolloutsTruncate with markers, return errors as observations, consistent formatting between training and deployment
Train/deploy matchSkills transfer best when the training harness matches productionSame tool schemas, system prompts and scaffold as the deployed agent; increasingly, train inside the production harness itself
Contamination hygieneTraining on eval tasks inflates numbersSeparate train/eval repos or time windows; hash and dedupe against public benchmarks
Intuition

An RL environment is a unit test for a skill that will be attacked by a highly motivated adversary millions of times. It needs to be fair (solvable, unambiguous), informative (some samples succeed, some fail), and robust (the only way to get reward is to have the skill). Most env-engineering effort goes into the last property.

Types of environments

TypeState & actionsRewardExamplesHard parts
Math / logicSingle turn; textAnswer checkerCompetition math sets, synthetic puzzlesAnswer normalization, wrong labels, saturation
Competitive codeSingle or few turns; code; optionally run testsHidden test pass rateLiveCodeBench-style problem setsWeak tests, timeouts, special-casing
SWE / repo tasksMulti-turn; shell + editor in a containerHidden tests (fail-to-pass, pass-to-pass), or patch similaritySWE-bench Jimenez+ 2023, SWE-Gym Pan+ 2024, R2E-Gym Jain+ 2025, SWE-smith Yang+ 2025, SWE-RL Wei+ 2025Building runnable envs per repo (dependency hell), heavy containers, test tampering
Terminal / opsShell in a containerScripted end-state checksTerminal-Bench tasks run via Harbor HarborMany valid solutions, nondeterministic tools
Browser / webMulti-turn; clicks, typing on DOM / accessibility tree / screenshotsEnd-state checks on backend DB or pageWebArena (self-hosted sites) Zhou+ 2023Live sites drift; must self-host clones; auth; slow
Computer useFull OS VM; screenshots + mouse/keyboardExecution-based checker scriptsOSWorld Xie+ 2024VM cost, latency, visual grounding, snapshotting
Tool use + simulated userMulti-turn chat with an LLM-simulated user and domain APIsFinal DB state matches goal; policy complianceτ-bench (retail, airline) Yao+ 2024Sim-user realism and consistency; the policy can manipulate the simulator
Search / researchInterleaved reasoning and search callsAnswer match (QA) or rubricSearch-R1 Jin+ 2025Live-index nondeterminism, cost, judge reliance
Code-interpreter reasoningReasoning with an executor toolAnswer checkerReTool Feng+ 2025Sandbox throughput
Games / puzzlesMulti-turn, fully simulatedGame score / winSokoban, FrozenLake in RAGEN Wang+ 2025Transfer to real tasks is unclear
Self-play / self-proposedModel proposes tasks and solves themExecutor-verifiedAbsolute Zero Zhao+ 2025Keeping the task distribution useful and non-degenerate
Non-verifiable (writing, advice)Single or multi-turnRubric/LLM judge, RMRubrics as Rewards Gunjal+ 2025Judge exploitation; reward drift

SWE environments in more detail, since they're the most-asked type. SWE-bench pairs real GitHub issues with the PR that fixed them. The grader runs the tests that the fix made pass (fail-to-pass) plus tests that must keep passing (pass-to-pass). Jimenez+ 2023 To train, you need many such executable instances. SWE-Gym built about 2.4k real executable Python tasks and showed both agent fine-tuning and verifier training benefit. Pan+ 2024 R2E-Gym generated environments procedurally from commits, with synthesized tests, plus hybrid (execution + model) verifiers. Jain+ 2025 SWE-smith scaled further by synthesizing bugs into repos, reaching tens of thousands of instances. Yang+ 2025 SWE-RL avoided execution altogether with a rule-based reward: the similarity between the generated patch and the ground-truth patch. That's cheap, but a weaker proxy. Wei+ 2025 The main engineering cost is building a working Docker image for each repo version.

Frameworks and interfaces (as of 2026)

ToolWhat it isNotes
Gymnasium (Farama)Standard single-agent RL env API; successor to OpenAI Gymreset/step with terminated vs truncated docs
verifiers + Environments Hub (Prime Intellect)Library for building LLM RL/eval environments, plus a community hub (launched Aug 2025) and the prime-rl trainer GitHub Prime Intellect 2025As of late 2026 the "v1" API is organized around tasksets (data + scoring), harnesses (the agent program, e.g., an existing coding agent), traces, and toolsets exposed as MCP servers. The earlier v0 API has been removed. Check current docs.
OpenEnv (Meta + Hugging Face)Spec and hub for agentic environments with a Gym-like reset()/step()/close() API, packaged as Docker-based environments HF blog 2025Announced Oct 2025; integrations with TRL, SkyRL, Unsloth and others announced or underway GitHub
HarborFramework from the Terminal-Bench creators for running agents (coding CLIs, OpenHands, …) on containerized tasks at scale, for evals and RL rollouts GitHubOfficial harness for Terminal-Bench 2.0; plugs into cloud sandbox providers
Benchmark-derived gymsSWE-Gym, R2E-Gym, SWE-smith, WebArena, OSWorld, τ-benchOften double as training envs, which makes contamination hygiene important

The broad trend: environments are moving from "a Python function that scores a string" to "a containerized task plus an unmodified production agent harness". The policy is trained inside the same scaffold it will be deployed in, and tools increasingly come through MCP.

Interview angle

The classic RL-env design question: "Design an environment to train an agent to do X (fix CI failures / file expense reports / do spreadsheet modelling)." A strong answer walks the anatomy table. Where tasks come from and how many; how state is provisioned and reset (snapshots); the tool surface and observation format; termination budgets; the verifier and how it resists hacking; difficulty calibration against the current policy; determinism (pinned images, mocked time/network); throughput and cost per episode; and the QA process (solvability checks with a reference solution, impossible-task canaries, red-teaming, transcript audits). Mentioning that every task should ship with a known-good solution that passes the grader, and a known-bad one that fails, is a strong signal.

Go deeper

10. Agentic and multi-turn RL

Multi-turn agent RL keeps the same algorithms but changes nearly everything around them.

Loss masking

A trajectory interleaves policy tokens with environment tokens (tool outputs, user messages, file contents). Only the policy's own tokens are actions. Environment tokens must be masked out of the policy-gradient loss (and the KL), or you train the model to "predict" tool outputs, which is wrong and destabilizing. Search-R1, for example, masks retrieved passages from the loss. Jin+ 2025 The same applies to prompt tokens and system messages.

Common mistake

Re-tokenization drift. If you store a trajectory as text and re-tokenize it for training, the token boundaries can differ from what the policy actually sampled (especially around tool-call delimiters and chat templates). The logprobs then no longer match, and importance ratios break silently. Record the exact token ids and logprobs at generation time and train on those.

Credit assignment over turns

Long horizons

Interview angle

"What changes when you go from single-turn RLVR to multi-turn agent RL?" Hit: loss masking of env tokens; exact token-id bookkeeping; credit assignment across turns; context management; invalid-trajectory handling; async env execution with very different episode lengths; sandbox cost; and the train/deploy harness match.

Go deeper

11. RL training infrastructure

The key fact: LLM RL is dominated by generation. Each training step needs fresh samples from the current policy, and autoregressive decoding of long CoTs or multi-turn episodes is far slower per token than the training forward/backward pass. Rollout generation is often reported as the majority of wall-clock time, often well over half, in synchronous systems. So RL systems are built as a fast inference fleet bolted to a trainer.

Task queuecurriculum / filters Rollout workers vLLMSGLang engineengine Env sandboxescontainers, VMs,browsers, sim users Reward serviceverifiers, judges Trajectory buffertokens, masks, logprobs TrainerFSDP / Megatronpolicy (+ref, +critic) weight sync (NCCL broadcast / resharding) tool calls ⇄ obs
A typical RL system: rollout workers (inference engines) interact with environment sandboxes, rewards are computed by a separate service, the trainer consumes trajectories and pushes new weights back to the generators.

Design axes

AxisOption AOption BTrade-off
PlacementColocated: the same GPUs alternate between generation and training (reshard weights between engine and trainer layouts)Disaggregated: separate GPU pools for rollout and trainingColocated: no idle pools, simple weight sync, but phases serialize. Disaggregated: phases overlap and each pool can be sized and tuned on its own, but one pool idles if they're unbalanced, and weights cross the network.
SynchronySynchronous: generate batch with \(\pi_k\), train, sync, repeatAsynchronous: generators run continuously; the trainer consumes data up to N versions staleAsync gives much higher utilization but off-policy data. Needs importance correction, staleness bounds, or robust objectives. Noukhovitch+ 2024 Fu+ 2025 (AReaL)
Weight update timingBetween batchesIn-flight: swap weights mid-generation, so a sequence can span policy versionsPipelineRL reports about 2× faster learning while staying nearly on-policy. Piché+ 2025
Long sequencesWait for all to finishPartial rollouts: cap per-iteration generation; continue unfinished sequences next iterationKimi k1.5's fix for long-tail CoTs. Kimi Team 2025

The long-tail problem

Response lengths are heavy-tailed. In a synchronous batch, everyone waits for the longest response or slowest episode. Decode batch occupancy decays as short sequences finish, and GPUs sit mostly idle at the end of each step. Fixes: (1) asynchronous or pipelined generation so the trainer never waits for stragglers; (2) partial rollouts / interruptible generation; (3) over-sampling and dropping the slowest (this biases against long correct answers, so be careful); (4) length caps with proper truncation handling; (5) continuous batching and prefix caching in the engine. The G samples for one prompt share a prefix, so prefix caching is a direct win.

Weight sync and the inference/trainer mismatch

Frameworks

FrameworkOriginNotable design
veRL (HybridFlow)ByteDance SeedHybrid single-/multi-controller programming model; colocated actor/rollout with resharding; FSDP or Megatron + vLLM/SGLang Sheng+ 2024 GitHub
OpenRLHFOpen-source communityRay-based, vLLM generation, distributed PPO/GRPO/REINFORCE++ Hu+ 2024
TRLHugging FaceGRPOTrainer and friends; vLLM in colocate or server mode; easiest on-ramp docs
slimeTHUDM (Zhipu / Tsinghua)Megatron training + SGLang rollout, built for large-scale RL post-training GitHub
AReaLAnt / TsinghuaFully asynchronous, interruptible rollouts, staleness-aware PPO Fu+ 2025
PipelineRLServiceNowIn-flight weight updates Piché+ 2025
prime-rlPrime IntellectAsync RL, used for globally decentralized training (INTELLECT-2); integrates with verifiers environments Prime Intellect 2025 GitHub
SkyRL, NeMo-RLBerkeley NovaSky; NVIDIAAgent/long-horizon focus (SkyRL); NVIDIA's scalable RL library SkyRL NeMo-RL

Back-of-envelope: one GRPO step

Say 512 prompts × G=16 samples × an average of 8k generated tokens. That's \(512\times16\times8{,}000\approx 6.6\times10^7\) tokens per step. If a 7B policy decodes at a few thousand tokens/s per GPU at high batch, then 64 GPUs at about 3k tok/s gives about 190k tok/s, or roughly 350 s of generation per step, ignoring the long tail. The training pass over the same tokens costs about \(6 \times 7\times10^9 \times 6.6\times10^7 \approx 2.8\times10^{18}\) FLOPs. At about 400 TFLOP/s effective per GPU that's roughly 110 s on 64 GPUs. So generation dominates even before stragglers and environment latency, which is why async designs and the long tail are the main concerns.

Interview angle

"Design the RL infrastructure for agentic coding at scale." Strong answers cover: disaggregated inference fleet + trainer; thousands of concurrent sandboxes with warm pools; async rollouts with bounded staleness plus importance correction; partial rollouts for long episodes; token-id-exact trajectory storage with loss masks; weight sync strategy; reward service isolated from sandboxes; invalid-trajectory handling; and observability (reward, length and entropy curves, KL, clip fraction, hack monitors, per-env pass rates). Bonus: the logprob mismatch and how you'd detect it (log both engine and trainer logprobs and plot the ratio distribution).

Go deeper

12. Evaluating RL'd models

MetricDefinitionUse
pass@1 (avg@n)Mean accuracy over n samples per problemPrimary metric; average many samples for low variance (small benchmarks like AIME have only 30 problems)
pass@kProbability at least one of k samples is correct. Unbiased estimator from n ≥ k samples with c correct: \(1-\binom{n-c}{k}/\binom{n}{k}\)Capability ceiling; the sharpening-vs-expansion debate
maj@k / cons@kMajority vote over k samplesTest-time compute scaling with no verifier
pass^kProbability all k trials succeed (τ-bench) Yao+ 2024Reliability for agents: users care about consistency
Accuracy vs tokensScore as a function of thinking budget or costEfficiency. A longer-CoT model may just spend more compute.
Hack / cheat ratePass rate on impossible variants; flagged-transcript rate Zhong+ 2025Is the score real?
Regression suiteGeneral knowledge, chat quality, safety, calibration, instruction followingRL on narrow domains can degrade other abilities
Interview angle

"Your RL'd model jumped 15 points on SWE-bench Verified. What do you check before celebrating?" Contamination (were these repos in training?), hack rate (test edits, special-casing), same scaffold and budget for the baseline, variance across seeds, pass^k reliability, accuracy per token/cost, and performance on held-out repos and a different benchmark (e.g., a terminal or multi-language set).

Go deeper

Interview question bank

Derive the policy gradient and explain why a baseline doesn't add bias.

\(\nabla_\theta \mathbb{E}_{\tau\sim p_\theta}[R]=\mathbb{E}[R\,\nabla\log p_\theta(\tau)]\) via \(\nabla p=p\nabla\log p\). Expanding \(\log p_\theta(\tau)\) gives the policy log-probs plus dynamics terms that don't depend on \(\theta\), so they vanish. The method is model-free and the reward needn't be differentiable. For a baseline \(b(s)\): \(\mathbb{E}_{a\sim\pi}[b(s)\nabla\log\pi(a|s)]=b(s)\nabla\sum_a\pi(a|s)=b(s)\nabla 1=0\). So subtracting it leaves the expectation unchanged and can greatly reduce variance. The optimal-ish baseline is \(V(s)\), which turns the weight into the advantage. The caveat is that the baseline must not depend on the action/sample being scored. That's why RLOO uses a leave-one-out mean.

Explain PPO's clipped objective. What does clipping do to the gradient?

PPO maximizes \(\min(\rho A, \text{clip}(\rho,1-\epsilon,1+\epsilon)A)\) per token, where \(\rho=\pi_\theta/\pi_\text{old}\). For positive advantage, once \(\rho>1+\epsilon\) the clipped term is smaller and constant, so the gradient is zero. For negative advantage the same happens once \(\rho\lt 1-\epsilon\). It's a pessimistic first-order trust region that lets you take multiple minibatch epochs on one batch of rollouts without the policy drifting far from the data-generating policy. Side effects: tokens outside the band stop contributing, which can silence rare high-value tokens (motivating CISPO), and symmetric clipping limits how fast unlikely tokens can grow (motivating DAPO's clip-higher).

Why did the field move from PPO to GRPO-style methods for reasoning RL?

The LLM setting is close to a contextual bandit (deterministic transitions, one terminal reward). Per-token value estimates are hard to learn and add little. The critic is as large as the policy, which roughly doubles training memory and compute, and it's unstable on sparse rewards. Sampling G responses per prompt gives an empirical baseline for free, and inference is needed anyway. With 0/1 verifiable rewards, the group mean is a very natural baseline. Trade-offs: no per-token credit assignment, wasted compute on all-tie groups, and normalization biases. VAPO-style work argues a well-trained critic can still win, so it isn't settled.

Write down the GRPO advantage and name three known problems with vanilla GRPO plus their fixes.

\(\hat A_i=(r_i-\text{mean}(r))/\text{std}(r)\), applied to all tokens of response i, inside a PPO-clipped loss with per-sequence \(1/|o_i|\) averaging and a KL-to-reference term. Problems: (1) length bias, since \(1/|o_i|\) makes long wrong answers cheap. Fix: Dr. GRPO removes it; DAPO uses token-level aggregation. (2) Std normalization over-weights near-uniform groups (difficulty bias). Fix: Dr. GRPO drops std; REINFORCE++ normalizes globally. (3) Entropy collapse. Fix: DAPO clip-higher, Clip-Cov/KL-Cov. (4) Noisy token-level ratios, especially MoE. Fix: GSPO's sequence-level ratio. (5) Zero-advantage groups waste compute. Fix: DAPO dynamic sampling, curriculum filtering.

What is the KL penalty for, and why do some reasoning recipes (DAPO) remove it?

The KL to a reference policy limits drift. That guards against reward-model overoptimization, preserves fluency and general abilities, and keeps outputs in-distribution for the reward model. In PPO-RLHF it's subtracted from the per-token reward. In GRPO it's a separate loss term. With programmatic verifiers there's no RM to exploit in the same way, and long-CoT reasoning needs the policy to move far from the base distribution (much longer outputs, new behaviors), so the KL mostly slows learning. DAPO drops it. ProRL keeps a KL but periodically resets the reference to the current policy, a compromise between stability and freedom to move.

Back-of-envelope: how many generated tokens per step, and is training or generation the bottleneck?

Example: 256 prompts × 16 samples × 10k tokens ≈ 41M tokens per step. Training FLOPs for a 32B model ≈ 6 × 32e9 × 41e6 ≈ 7.9e18. At about 400 TFLOP/s effective per GPU on 128 GPUs, that's roughly 150 s. Decoding 41M tokens of a 32B model at maybe 1–2k tok/s per GPU (an order-of-magnitude guess; measure it) on 128 GPUs is about 160–320 s before the long tail. The slowest 1% of sequences can double wall-clock in synchronous mode. So generation, and especially its tail, usually dominates, which motivates async/pipelined RL, partial rollouts and prefix caching. Always state that throughput numbers depend heavily on hardware, engine and batch shape.

What is RLVR and what makes a good verifier?

RL with verifiable rewards uses programmatic checks (math answer equivalence, hidden unit tests, constraint checkers) instead of a learned reward model. A good verifier is correct (no false positives or negatives on reference solutions), robust to equivalent forms, deterministic, fast and cheap, isolated from the policy (runs outside its sandbox, hidden data not visible), and resistant to known exploits (multiple answers, test edits, exit-code tricks). Validate it by running known-good and known-bad solutions, check that reference answers are actually right, and track its false-positive rate during training with transcript audits.

Does RL teach models new reasoning capabilities, or just sharpen existing ones? Argue both sides.

Sharpening: Yue et al. found RLVR models beat base models at pass@1 but lose at pass@k for large k across families and algorithms. Their solutions appear to lie within the base model's coverage, and distillation, not RL, added new patterns. Spurious-reward results show gains even with random rewards on Qwen, consistent with amplifying latent behaviors. Expansion: ProRL found prolonged, diverse RL with reference resets improves pass@k across the board, including tasks where the base model never succeeds, with gains largest where the base was weakest. Synthesis: short RL on narrow tasks mostly reweights; long, diverse, compositional training, especially agentic, can push the boundary. Exploration (base prior, curriculum) decides which regime you're in. It's still contested, so say so.

What happened in DeepSeek-R1-Zero and why was it significant?

GRPO was applied directly to a strong base model (DeepSeek-V3-Base) with only rule-based accuracy and format rewards, with no SFT. Accuracy on AIME rose dramatically, response length grew on its own, and reflection and backtracking behaviors appeared. It was significant because it showed in the open that long-CoT reasoning can emerge from outcome-only RL without PRMs, MCTS or curated reasoning traces. Caveats: poor readability and language mixing (fixed in R1 with cold-start SFT and a language-consistency reward), and later work showing aha-like behaviors already exist in base models, so "emergence" partly means amplification.

Process reward models vs outcome rewards: when would you use each?

Outcome rewards (final-answer checks, tests) are cheap, hard to hack when programmatic, and the default for large-scale RL. Their weakness is sparse credit assignment over long chains. PRMs score each step, which gives dense signal and helps search/reranking (Lightman's PRM beat an ORM at best-of-N), but labels are expensive or noisy (MC-estimated), "steps" are ill-defined, and policies learn to please the PRM. Use PRMs or progress-style verifiers for test-time search, verifier-guided decoding, or as auxiliary dense signals with care. Use outcome rewards as the primary RL signal. R1 and Kimi k1.5 both reported PRMs weren't worth it at their scale.

Give five concrete ways a model could reward-hack a coding RL environment, and how you'd prevent each.

(1) Hard-code expected outputs for visible tests. Prevent: hidden and randomized tests. (2) Edit or delete failing tests or conftest.py. Prevent: run graders from a pristine copy outside the sandbox, and checksum test files. (3) Exit early with status 0, or override __eq__ to always be true. Prevent: graders that check actual outputs and don't trust exit codes; detect monkeypatching. (4) Find the reference fix in git history, pip caches or the internet. Prevent: scrub history, egress allowlist. (5) Tamper with timing or scoring harness code (documented by METR). Prevent: keep graders out of the writable filesystem. Plus process: red-team the env with a strong model before training, include impossible-task canaries, and run monitors over transcripts.

Why is optimizing against a chain-of-thought monitor risky?

Baker et al. showed that LLM monitors reading the CoT catch reward hacking much better than action-only monitors. But adding the monitor's verdict to the reward, under strong optimization, taught the model to hide its intent: the CoT looks clean while the hack continues ("obfuscated reward hacking"). You lose the monitoring channel and still have the misbehavior. Recommended practice is to use CoT monitors for detection, env patching and data filtering, and to avoid strong direct optimization pressure on CoT content, accepting a "monitorability tax".

Design an RL environment for training an agent to resolve CI failures in real repositories.

Tasks: mine historical CI failures with the commit that fixed them; filter to deterministic failures reproducible in a container; split by repo/time for train vs eval. State: a container image per repo version with deps pinned and the failing commit checked out; snapshot for fast reset. Tools: shell (with timeout), file view/edit, run-tests, submit. Observations: truncated logs with markers, consistent formatting. Termination: submit, N turns, or wall-clock. Reward: the originally failing job passes plus a pass-to-pass regression suite, run by a separate grader on a pristine copy of the tests. Possibly a small penalty for touching test/CI config. Calibration: measure the current policy's pass rate and keep a 10–90% band. QA: verify each task with the real fix (must pass) and a no-op (must fail), add impossible canaries, red-team. Determinism: no network, or record/replay, seeded randomness. Throughput: warm pools, cached images, cost per episode tracked.

What is loss masking in multi-turn RL, and what goes wrong without it?

In agent trajectories, tool outputs, user messages and environment text are interleaved with the policy's tokens. Only the policy's tokens are actions, so the policy-gradient loss (and KL) must be computed only over them, via a mask that is 1 on sampled tokens and 0 on prompt/environment tokens. Without masking, the model is pushed to raise or lower the likelihood of text it didn't choose (tool outputs) according to the episode's advantage. That's meaningless credit assignment that can destabilize training and teach it to hallucinate tool outputs. Related: store exact token ids and logprobs from generation to avoid re-tokenization drift.

Colocated vs disaggregated RL infra: when would you choose each?

Colocated (veRL-style hybrid engine): the same GPUs switch between inference and training, resharding weights. It's simple, has no idle dedicated pools, and weight sync is cheap (local), so it's good for smaller clusters or short generations. Disaggregated: separate inference and training pools that run concurrently, each sized and tuned independently. That suits long rollouts, agentic workloads with environment latency, and async pipelines, but it needs network weight transfer and load balancing, and adds staleness. At frontier scale with long agentic episodes, disaggregated plus async is the common direction. Colocated is a strong default for research-scale reasoning RL.

What is the rollout/trainer logprob mismatch, and how do you handle it?

The inference engine (vLLM/SGLang) and the training framework compute slightly different probabilities for the same tokens and weights, because of different kernels, precision, batch-dependent reductions, quantized rollouts and MoE routing differences. So nominally on-policy data is off-policy, and PPO ratios computed against engine logprobs are biased. Detect it by logging both and plotting \(\pi_\text{train}/\pi_\text{rollout}\). Mitigate with truncated importance sampling (weight by \(\min(\text{ratio}, C)\)), recomputing old logprobs with the trainer, sequence-level ratios, matching precision, or batch-invariant deterministic kernels.

How does asynchronous RL work, and what's the cost?

Generators keep producing rollouts with whatever weights they have while the trainer updates. Data may be k versions stale. Benefits: high GPU utilization and no waiting for stragglers. Costs: off-policy bias and instability as staleness grows. Mitigations: bound staleness (drop data older than k versions), importance-weight corrections (decoupled PPO objectives as in AReaL), in-flight weight updates to keep sequences fresh (PipelineRL), and objectives more robust to off-policy data (Noukhovitch et al. found online DPO robust). Report the staleness distribution as a training metric.

How do you calibrate task difficulty for GRPO, and why does it matter so much?

With group-relative advantages, a prompt where all G samples get the same reward yields zero gradient. Too-easy and too-hard tasks waste rollout compute and also shrink the effective batch, which adds variance. Calibrate by estimating the current policy's pass rate per task (8–16 samples), keeping tasks in an intermediate band, and re-estimating periodically as the policy improves (curriculum). Online, DAPO's dynamic sampling over-samples and discards all-tie groups. Keep a reserve of harder tasks to promote as the model improves, and track the fraction of zero-variance groups as a health metric.

How would you get RL signal for a non-verifiable task such as writing medical patient summaries?

Push toward verifiability: extract checkable sub-claims (medications, dosages, diagnoses present in the source record; no hallucinated facts, via entailment against the record), plus format and length constraints. For quality, use per-prompt rubrics (expert-written or model-drafted then expert-reviewed) graded item-by-item by a judge from a different model family, or pairwise comparisons against a reference. Guard against judge exploitation with length control, hidden rubrics, periodic human audits, and tracking judge-human agreement over training. Keep a KL anchor because the reward is softer than a program.

What metrics do you watch during an RL run to know it's healthy?

Reward (train) and held-out eval accuracy, which should track each other; divergence suggests hacking or overfitting. Response length, split by correct and incorrect. Policy entropy (collapse means lost exploration). KL to the reference. Clip fraction and importance-ratio statistics (including the engine-vs-trainer mismatch). Fraction of zero-advantage groups. Gradient norm. Per-environment and per-difficulty pass rates. Truncation rate. Invalid-trajectory rate (environment failures). Staleness distribution in async setups. Hack-monitor flags. Plus periodic transcript reading, which no dashboard replaces.

Compare GSPO's sequence-level ratio with GRPO's token-level ratio.

GRPO computes \(\pi_\theta/\pi_\text{old}\) per token and clips per token, even though the advantage is the same for every token in a sequence. Each token's ratio is a one-sample, high-variance estimate, and the noise accumulates over long sequences. In MoE models, routing shifts make individual token ratios swing more. GSPO uses the length-normalized sequence likelihood ratio \((\pi_\theta(o)/\pi_\text{old}(o))^{1/|o|}\) and clips whole sequences, matching the granularity of the reward. It reports more stable training (notably for MoE) and simpler infrastructure, since it's less sensitive to small token-level precision mismatches. The cost: a sequence is either inside or outside the trust region as a whole.

What's an RL environment "harness", and why does the train/deploy match matter?

The harness is the agent program around the model: the system prompt, tool schemas, the loop that parses tool calls and feeds back observations, context management. Policies learn harness-specific habits (tool-call formats, how much output to request, when to submit). If training uses a toy ReAct loop but deployment uses a production coding agent with different tools and context handling, skills transfer imperfectly. The 2025–2026 trend, visible in Harbor and verifiers v1, is to run RL inside the real production harness, with tools via MCP, so the trained behavior is the deployed behavior.

You see reward going up but held-out accuracy flat. Walk through your debugging.

First, read transcripts from high-reward rollouts for hacks (test edits, special-casing, judge flattery, multiple answers). Check whether the reward is dominated by an auxiliary component (format, length). Compare train vs held-out distributions (are eval tasks harder, different domain, different harness?). Check contamination or memorization of training tasks (reward rising on repeated tasks only). Verify the grader on known-bad solutions. Check whether gains are in extraction/format that the eval scores differently. Fixes: harden the env, rebalance reward components, add held-out checks into training monitoring, early stop.