RL for LLMs & RL Environments
Reinforcement learning went from the last, small step of RLHF to the main driver of frontier capability gains. Reasoning models, coding agents and computer-use agents all get their edge from large-scale RL against verifiable rewards inside carefully built environments. This page covers the RL you need (policy gradients, PPO, GRPO and its fixes), how reasoning models are trained, why reward hacking is the central failure mode, and, in the most depth, how RL environments are designed, built and scaled. Frontier labs and the fast-growing set of RL-environment startups now interview directly on that last topic.
TL;DR: the things to be able to say out loud
- Policy gradient: \(\nabla J = \mathbb{E}[\nabla \log \pi_\theta(a|s)\,A(s,a)]\). Raise the log-probability of actions that did better than expected. Everything else on this page is variance reduction, stability and systems engineering built on that.
- PPO = policy gradient + importance ratio + clipping (a cheap trust region) + a learned value function for advantages (GAE). Classic RLHF adds a KL penalty to a frozen reference model.
- LLM as policy: state = prompt plus tokens so far, action = next token, reward is usually given once at the end of the sequence. Single-turn tasks are effectively contextual bandits over whole responses. Agent tasks are real multi-turn MDPs.
- GRPO drops the critic. It samples G responses per prompt and uses \((r_i-\text{mean})/\text{std}\) as each response's advantage. It's cheap and works well with 0/1 verifiable rewards.
- GRPO's known flaws and their fixes: length-normalization bias (Dr. GRPO, DAPO token-level loss), entropy collapse (DAPO clip-higher, Clip-Cov/KL-Cov), zero-signal groups (DAPO dynamic sampling), noisy token-level ratios especially in MoE (GSPO sequence-level ratio, CISPO).
- RLVR: rewards from programs (answer checkers, unit tests, format checks) instead of a learned reward model. Much harder to hack than a reward model, though not impossible. DeepSeek-R1-Zero showed long chain-of-thought emerging from pure RLVR on a strong base model.
- The open debate: does RL teach new capabilities or just sharpen what's already there (better pass@1, sometimes worse pass@k at large k)? The evidence points both ways, and the answer depends on training length, task diversity and the base model.
- Reward hacking is the main failure mode: special-casing tests, deleting tests, sycophancy, gaming LLM judges. Mitigations: robust graders, held-out checks, monitors, KL/regularization, environment hardening and auditing rollouts.
- An RL environment = task distribution + initial state + action/tool interface + observation formatting + transition dynamics (often a sandbox) + termination rule + reward/verifier. A good one is verifiable, hard to hack, calibrated in difficulty, diverse, reproducible and cheap to run in parallel.
- Infrastructure: RL for LLMs is mostly generation. Rollouts run on vLLM/SGLang, the trainer runs on FSDP/Megatron, and updated weights are synced back to the generators. Async or pipelined designs trade off-policy staleness for throughput. Long-tail rollouts are the bottleneck.
- Multi-turn agent RL: mask tool/environment output tokens out of the loss, assign credit across turns, and budget for long horizons, sandboxes and flaky tools.
1. RL fundamentals, compactly
You will be asked to define these crisply. Interviewers at labs often start here to check that you know RL itself and haven't only memorized "GRPO".
The MDP
A Markov decision process is a tuple \((\mathcal{S}, \mathcal{A}, P, R, \gamma, \rho_0)\): states, actions, transition dynamics \(P(s'|s,a)\), reward \(R(s,a)\), discount \(\gamma\in[0,1]\) and initial-state distribution \(\rho_0\). Markov means the next state depends only on the current state and action. A policy \(\pi_\theta(a|s)\) is a distribution over actions. A trajectory (episode, rollout) is \(\tau = (s_0,a_0,r_0,s_1,\dots)\). The return is the discounted sum of rewards \(G_t = \sum_{k\ge 0}\gamma^k r_{t+k}\). The objective is $$J(\theta) = \mathbb{E}_{\tau\sim\pi_\theta}\Big[\sum_t \gamma^t r_t\Big].$$ Sutton & Barto 2018
Value functions, advantage, Bellman
| Quantity | Definition | Meaning |
|---|---|---|
| State value \(V^\pi(s)\) | \(\mathbb{E}_\pi[G_t\mid s_t=s]\) | How good it is to be in \(s\) and then follow \(\pi\) |
| Action value \(Q^\pi(s,a)\) | \(\mathbb{E}_\pi[G_t\mid s_t=s,a_t=a]\) | How good it is to take \(a\) in \(s\) and then follow \(\pi\) |
| Advantage \(A^\pi(s,a)\) | \(Q^\pi(s,a)-V^\pi(s)\) | How much better \(a\) is than the policy's average action. Zero-mean under \(\pi\). |
| TD error \(\delta_t\) | \(r_t+\gamma V(s_{t+1})-V(s_t)\) | One-step estimate of the advantage (exact if \(V=V^\pi\), in expectation) |
The Bellman expectation equation writes value recursively: \(V^\pi(s)=\mathbb{E}_{a\sim\pi,s'\sim P}[r+\gamma V^\pi(s')]\). The Bellman optimality equation replaces the expectation over actions with a max: \(Q^*(s,a)=\mathbb{E}_{s'}[r+\gamma\max_{a'}Q^*(s',a')]\). Value-based methods (Q-learning, DQN) iterate on the optimality equation. Policy-gradient methods optimize \(\pi\) directly and use value functions only to reduce variance. LLM RL is almost entirely in the policy-gradient family, because the action space (a vocabulary of about 100k+ tokens, applied autoregressively) makes a max over actions meaningless at the sequence level, and because we already start from a very good policy, the pretrained model.
On-policy vs off-policy, exploration, credit assignment
- On-policy methods (REINFORCE, PPO, GRPO) learn from data sampled by the current policy. Every update needs fresh samples. Off-policy methods (Q-learning, SAC, DPO-style offline objectives) can learn from data generated by another policy, which means replay buffers or static datasets. In practice LLM RL is "nearly on-policy". PPO/GRPO take a few gradient steps per batch, and asynchronous systems let data go a few versions stale. Importance ratios \(\pi_\theta/\pi_{\text{old}}\) correct for the small mismatch.
- Exploration: the policy only learns from what it samples. If a correct solution has probability \(10^{-6}\), you will essentially never see the reward. With LLMs, exploration is shaped by the base model's prior, by sampling temperature, by group size, and by curriculum (tasks just hard enough that some samples succeed). Entropy collapse, where the policy becomes deterministic too early, is the LLM version of premature exploitation.
- Credit assignment: a single 0/1 reward at the end of a 10,000-token chain of thought must somehow be attributed to the tokens that mattered. Outcome-reward RL spreads it uniformly (every token in a successful response gets the same advantage) and relies on averaging over many samples to separate signal from noise. Process rewards, value functions and turn-level rewards try to do better.
Supervised learning says "make this output more likely." Policy-gradient RL says "make whatever you just did more likely, in proportion to how much better than expected it turned out." The model generates its own training data, and the reward decides the sign and size of each update. That's why RL can go beyond the demonstrations (it can find solutions nobody wrote down) and also why it can find exploits nobody intended.
"What's the difference between \(V\), \(Q\) and \(A\), and why do we use \(A\) in the gradient?" A strong answer: \(A\) is centered, so subtracting \(V\) removes a large state-dependent offset that adds variance without changing the expected gradient. Also expect to say why LLM RL is on-policy-ish, and what credit assignment means when the reward is a single number per response.
- Sutton & Barto, Reinforcement Learning: An Introduction (free online): the canonical textbook; chapters 3, 6 and 13 cover this section.
- OpenAI Spinning Up: Key Concepts in RL: concise, deep-learning-oriented notation.
- Nathan Lambert, RLHF Book: bridges classical RL to LLM post-training.
2. Policy gradients: REINFORCE to GAE
The log-derivative trick
We want \(\nabla_\theta \mathbb{E}_{\tau\sim\pi_\theta}[R(\tau)]\), but the expectation is over a distribution that itself depends on \(\theta\), and the reward is not differentiable (it's a unit-test result). The trick is \(\nabla p_\theta = p_\theta \nabla\log p_\theta\), so
$$\nabla_\theta J = \mathbb{E}_{\tau\sim\pi_\theta}\big[R(\tau)\,\nabla_\theta\log p_\theta(\tau)\big],\qquad \log p_\theta(\tau)=\sum_t\log\pi_\theta(a_t|s_t) + \underbrace{\textstyle\sum_t \log P(s_{t+1}|s_t,a_t)}_{\text{no }\theta}.$$
The dynamics terms drop out. You need neither a model of the environment nor a differentiable reward, only the ability to sample and score. This is the REINFORCE estimator (Williams, 1992). In code, it is just a weighted negative log-likelihood: loss = -(advantage.detach() * logprobs).mean().
Baselines and variance reduction
The raw estimator is unbiased but has very high variance. Three standard fixes:
- Causality / reward-to-go: an action can't affect past rewards, so weight \(\nabla\log\pi(a_t|s_t)\) by \(G_t\), not the full return. (With a single terminal reward, as in most LLM RL, these coincide.)
- Baselines: subtracting any \(b(s_t)\) that doesn't depend on the action leaves the gradient unbiased, because \(\mathbb{E}_{a\sim\pi}[\nabla\log\pi(a|s)]=\nabla\sum_a\pi(a|s)=\nabla 1=0\). The best common choice is \(b=V(s)\), which gives the advantage. In LLM RL the baseline is often empirical: the mean reward of other samples for the same prompt (RLOO, GRPO), or a batch mean (REINFORCE++).
- Normalization: whitening advantages (divide by std) stabilizes the step size across prompts and reward scales. It also introduces subtle biases (§5).
Generalized Advantage Estimation (GAE)
With a learned critic \(V_\phi\), you can estimate advantages anywhere between the one-step TD error (low variance, biased if \(V\) is wrong) and the Monte-Carlo return minus baseline (unbiased, high variance). GAE takes an exponentially weighted mix: $$\hat A_t^{\text{GAE}(\gamma,\lambda)}=\sum_{l\ge 0}(\gamma\lambda)^l\,\delta_{t+l},\qquad \delta_t=r_t+\gamma V_\phi(s_{t+1})-V_\phi(s_t).$$ \(\lambda=0\) gives TD(0). \(\lambda=1\) gives the MC return minus \(V\). Schulman+ 2015 In LLM RLHF with PPO, typical settings were \(\gamma=1\) and \(\lambda\) close to 1 (0.95–1.0), and the reward was sparse: the reward-model score at the last token, plus a per-token KL penalty.
Saying "the baseline makes the gradient biased." It doesn't, provided the baseline doesn't depend on the action being scored. A subtle exception matters in practice: if the baseline for sample \(i\) includes sample \(i\)'s own reward (as GRPO's group mean does), there is a small bias. That's exactly why RLOO uses the leave-one-out mean.
"Derive the policy gradient" is a common whiteboard ask. Hit three beats: (1) log-derivative trick, (2) dynamics terms vanish, so it's model-free, (3) a baseline keeps it unbiased because the expected score function is zero. Bonus: explain GAE's \(\lambda\) as a bias-variance dial and why LLM people often drop the critic altogether.
- Schulman et al., GAE (2015): the original GAE derivation.
- Spinning Up: Intro to Policy Optimization: clean derivation of the REINFORCE and baseline results.
3. PPO and classic RLHF
Why trust regions
Plain policy gradient takes one step per batch of samples. If you take several steps on the same batch, or one large step, the policy can move far from the one that generated the data, so the gradient estimate becomes wrong and performance collapses. TRPO enforced a hard KL constraint between old and new policy with a second-order method. Schulman+ 2015 PPO gets most of that benefit with a first-order clipped surrogate. Schulman+ 2017
The clipped objective
Let \(\rho_t(\theta)=\dfrac{\pi_\theta(a_t|s_t)}{\pi_{\theta_\text{old}}(a_t|s_t)}\) be the importance ratio. Then
$$L^{\text{CLIP}}(\theta)=\mathbb{E}_t\Big[\min\big(\rho_t\hat A_t,\ \text{clip}(\rho_t,1-\epsilon,1+\epsilon)\hat A_t\big)\Big],\qquad \epsilon\approx 0.2.$$
If \(\hat A_t>0\), the objective stops rewarding increases of \(\rho_t\) above \(1+\epsilon\). If \(\hat A_t\lt 0\), it stops rewarding decreases below \(1-\epsilon\). The min makes it a pessimistic bound: clipping can only remove incentive, never add it. Once a token's ratio leaves the band in the "already improved" direction, its gradient is exactly zero. That's the property later algorithms (DAPO, CISPO) tinker with.
The full PPO-RLHF setup
InstructGPT-style RLHF uses four models: the policy (trained), a frozen reference policy (the SFT model), a reward model (frozen, trained beforehand on human preference pairs with a Bradley-Terry loss), and a value model/critic (trained, often initialized from the reward model). Ouyang+ 2022 Christiano+ 2017 The per-token reward is $$r_t = \underbrace{r_\text{RM}(x,y)\cdot\mathbb{1}[t=T]}_{\text{sequence score at the end}} - \beta\,\log\frac{\pi_\theta(y_t|x,y_{\lt t})}{\pi_\text{ref}(y_t|x,y_{\lt t})}.$$ The KL term keeps the policy near the reference. It stops reward-model exploitation ("overoptimization": the proxy score rises while the true quality falls) and preserves fluency and diversity. Gao+ 2022 The value loss is a regression of \(V_\phi(s_t)\) onto the returns, often also clipped.
| PPO-RLHF component | Why it exists | Cost / pain |
|---|---|---|
| Policy \(\pi_\theta\) | The thing being trained | Full training state (weights, grads, Adam: about 16 bytes/param in mixed precision) |
| Reference \(\pi_\text{ref}\) | KL anchor against drift and reward hacking | A second forward pass per batch |
| Reward model | Turns preferences into a scalar | Hackable proxy; needs fresh preference data |
| Critic \(V_\phi\) | Per-token baselines via GAE | Same size as the policy, so it roughly doubles training memory. Hard to train well on sparse, end-of-sequence rewards. |
The critic is the most fragile part for LLMs. It has to predict, from a partial response, the expected final reward. That's a hard regression problem on top of a non-stationary policy. That fragility, plus the memory cost, is the main reason the field moved to critic-free methods that estimate the baseline by sampling several responses per prompt.
Expect: "Walk me through one PPO step for an LLM." Generate responses, score them with the RM, compute per-token KL vs the reference, compute values and GAE, then run several minibatch epochs of clipped policy loss + value loss. Probes: what does \(\beta\) trade off (reward vs. drift); what happens when \(\epsilon\) is too large (instability) or too small (slow learning); why is the KL in the reward in PPO-RLHF but in the loss in GRPO?
- PPO paper (2017): short and readable.
- InstructGPT (2022): the canonical three-stage SFT → RM → PPO pipeline.
- Scaling Laws for Reward Model Overoptimization: why the KL leash matters, quantified.
4. The LLM as a policy
Mapping the MDP onto a language model:
| MDP element | Single-turn (math, code, chat) | Multi-turn agent (SWE, browser, tools) |
|---|---|---|
| State \(s_t\) | Prompt + tokens generated so far | Full transcript: system prompt, task, prior actions, tool outputs, plus hidden environment state (filesystem, DB, web page) |
| Action \(a_t\) | Next token (or the whole response as one "action") | Token-level, but usually reasoned about per turn: one tool call or message |
| Transition | Deterministic: append the token | Stochastic/external: the environment runs the tool and returns an observation |
| Reward | Once, at the end (verifier or RM) | Usually at the end (tests pass, task completed); sometimes intermediate |
| Horizon | Hundreds to tens of thousands of tokens | Tens to hundreds of turns, potentially millions of tokens of context over the episode |
The bandit view. With a deterministic "append token" transition and a single terminal reward, single-turn RL is effectively a contextual bandit: context = prompt, arm = complete response, payoff = reward. That's why simple REINFORCE-style estimators with a per-prompt baseline work so well (RLOO's main argument). Ahmadian+ 2024 Per-token value functions buy less here than in Atari or robotics, because there are no intermediate rewards to bootstrap from.
The token-level view still matters for the implementation. The gradient is a sum over tokens, importance ratios and clipping are per token in PPO/GRPO, and the KL penalty is per token. Choices like how you average over tokens (per sequence or per batch) change the effective weighting of long vs short responses. Those choices turned out to matter a lot (§5).
The multi-turn view is a real MDP: the environment's responses are not under the policy's control and must be excluded from the loss. Rewards may arrive after many turns, and the "state" includes things the model can't see (a database, a remote browser). §10 covers this.
"Is token-level or sequence-level the right granularity for an action?" A strong answer: the reward lives at the sequence level, the optimization happens at the token level, and the mismatch between them causes real bugs (length bias, ratio noise). GSPO's argument is that the importance ratio should match the granularity of the reward. For agents, the turn is often the natural unit for credit assignment.
5. GRPO and its successors
GRPO: Group Relative Policy Optimization
Introduced in DeepSeekMath and made famous by DeepSeek-R1. Shao+ 2024 For each prompt \(q\), sample a group of \(G\) outputs \(\{o_1,\dots,o_G\}\) from \(\pi_{\theta_\text{old}}\), score them, and give every token of output \(i\) the same advantage: $$\hat A_{i,t} = \frac{r_i-\text{mean}(r_1,\dots,r_G)}{\text{std}(r_1,\dots,r_G)}.$$ The objective is a PPO-style clipped surrogate averaged per sequence, then over the group, with the KL term moved from the reward into the loss: $$J_\text{GRPO}(\theta)=\mathbb{E}\Bigg[\frac{1}{G}\sum_{i=1}^G\frac{1}{|o_i|}\sum_{t=1}^{|o_i|}\Big(\min\big(\rho_{i,t}\hat A_{i,t},\ \text{clip}(\rho_{i,t},1\!-\!\epsilon,1\!+\!\epsilon)\hat A_{i,t}\big)-\beta\,\mathbb{D}_\text{KL}[\pi_\theta\|\pi_\text{ref}]\Big)\Bigg]$$ where \(\rho_{i,t}=\pi_\theta(o_{i,t}|q,o_{i,\lt t})/\pi_{\theta_\text{old}}(o_{i,t}|q,o_{i,\lt t})\). The KL is estimated per token with the low-variance "k3" estimator \(\frac{\pi_\text{ref}}{\pi_\theta}-\log\frac{\pi_\text{ref}}{\pi_\theta}-1\).
Why it caught on: no critic (saving roughly a policy-sized model of memory and a hard training problem); the group mean is a good baseline for 0/1 verifiable rewards; and it's simple. Typical group sizes are 8–64 samples per prompt.
Key property: if all \(G\) samples get the same reward (all right or all wrong), every advantage is zero and the prompt contributes no gradient. Prompts that are too easy or too hard are wasted compute. That's the root of curriculum and difficulty-filtering practices (§9).
Known problems with vanilla GRPO
- Length bias from \(1/|o_i|\). Each sequence's loss is divided by its own length. For a negative advantage, a longer wrong answer gets a smaller per-token penalty, so the model learns that being wrong at length is cheaper. Response length inflates, especially for incorrect outputs. Liu+ 2025 (Dr. GRPO)
- Difficulty bias from std normalization. Dividing by the group std up-weights prompts where almost all samples agree (low std), meaning very easy or very hard ones. Dr. GRPO argues this skews the optimization.
- Entropy collapse. With symmetric clipping, low-probability "exploration" tokens with positive advantage can only rise by a factor of \(1+\epsilon\) per step, while already-likely tokens dominate. Entropy drops fast early in training and performance plateaus. One paper fits an empirical relationship between performance and entropy, \(R\approx -a e^{H}+b\), suggesting performance is "bought" with entropy. Cui+ 2025
- Token-level importance ratios are noisy. The ratio of a single token's probability under two nearby policies is a high-variance estimate when one sequence-level reward is spread over thousands of tokens. This gets worse with MoE models, where routing changes between \(\pi_\text{old}\) and \(\pi_\theta\) can swing individual token probabilities. Zheng+ 2025 (GSPO)
- Clipping as an unintended prior. One study found that GRPO with random rewards still improved Qwen2.5-Math-7B on MATH-500 by about 21 points (vs. about 29 with true rewards), and attributed it to a clipping bias that amplifies behaviors the base model already favors. The effect largely didn't transfer to Llama or OLMo. Shao+ 2025 Treat RLVR gains on a single model family with suspicion.
The successor family
RLOO (REINFORCE Leave-One-Out): sample \(k\) responses and use as baseline for each the mean reward of the other \(k-1\). This is unbiased, critic-free, with no clipping or ratio in the simplest form. The paper argued PPO's machinery is mostly unnecessary in the LLM bandit setting. Ahmadian+ 2024 (Mean-centering with the full group, as in GRPO without the std division, gives exactly \(\tfrac{k-1}{k}\) times the RLOO advantage, so the two differ only by a constant scale.)
REINFORCE++: critic-free, but normalizes advantages over the global batch instead of per prompt. The claim is that per-prompt normalization is biased and overfits, and global normalization is effectively unbiased as the batch grows. It has a variant with a group baseline for reasoning tasks. Hu+ 2025
Dr. GRPO ("GRPO Done Right"): removes the \(1/|o_i|\) per-sequence length normalization and the std division, which fixes the length and difficulty biases above. It reports similar accuracy with better token efficiency. Liu+ 2025
DAPO (ByteDance Seed / Tsinghua): four changes on top of GRPO, with no KL term, for long-CoT reasoning Yu+ 2025:
- Clip-Higher: decoupled bounds \(\epsilon_\text{low}=0.2\), \(\epsilon_\text{high}\approx 0.28\). Unlikely tokens can grow more, which counters entropy collapse.
- Dynamic Sampling: drop prompts whose groups are all-correct or all-wrong (zero advantage) and keep sampling until the batch is full of informative groups.
- Token-level policy gradient loss: average over all tokens in the batch instead of per sequence, so long sequences are weighted by their length and long degenerate outputs get properly penalized.
- Overlong reward shaping: a soft length penalty near the max length, and masking of truncated samples, so the reward isn't noisy when a good answer gets cut off.
GSPO (Qwen team): uses a sequence-level importance ratio, the length-normalized likelihood ratio \(s_i=\big(\pi_\theta(o_i|q)/\pi_{\theta_\text{old}}(o_i|q)\big)^{1/|o_i|}\), and clips whole sequences. The reward is per sequence, so the off-policy correction should be too. It reports better stability, especially for MoE RL, and credits it for improvements in Qwen3. Zheng+ 2025
CISPO (MiniMax-M1): instead of zeroing the gradient of clipped tokens, it clips the importance-sampling weight (as a stop-gradient coefficient) and keeps every token's gradient. Rare but important "fork" tokens, like a reflective "wait", aren't silenced. MiniMax 2025
VAPO: brings the value function back (value-pretraining, decoupled GAE \(\lambda\) for actor and critic, length-adaptive GAE) and argues a well-trained critic beats critic-free methods on long CoT. Yue+ 2025 Clip-Cov / KL-Cov restrict updates on the high-covariance tokens that drive entropy collapse. Cui+ 2025
| Method | Critic? | Advantage / baseline | Ratio & clipping | Loss aggregation | Main fix / claim |
|---|---|---|---|---|---|
| REINFORCE | No | Return − (optional) baseline | None (pure on-policy) | Any | Simplest unbiased estimator |
| PPO | Yes | GAE from learned \(V\) | Token ratio, symmetric clip ε≈0.2 | Token mean | Trust region cheaply; dense per-token baselines |
| RLOO | No | Leave-one-out group mean | Often none | Sequence | Bandit framing; PPO machinery unnecessary |
| GRPO | No | Group mean, ÷ group std | Token ratio, symmetric clip | Per-sequence mean, then group mean | No critic; good with 0/1 rewards |
| Dr. GRPO | No | Group mean (no std) | Token ratio, clip | No per-sequence length norm | Removes length and difficulty bias |
| DAPO | No | Group-normalized | Token ratio, asymmetric clip (clip-higher) | Token-level over batch | Entropy collapse, zero-signal groups, long-CoT stability; no KL |
| REINFORCE++ | No | Global batch normalization | Token ratio, clip | Token | Less biased normalization; robust in general RLHF |
| GSPO | No | Group-normalized | Sequence-level ratio (length-normalized), sequence clip | Sequence | Ratio matches reward granularity; MoE stability |
| CISPO | No | Group-normalized | Clip the IS weight (stop-grad), keep all token grads | Token | Rare tokens keep gradient; efficiency |
| VAPO | Yes | GAE with value-pretraining, length-adaptive λ | Token, clip-higher | Token | A good critic still wins on long CoT |
This algorithm zoo moves monthly. A large Meta study (over 400k GPU-hours) concluded that many of these choices (loss aggregation, normalization, off-policy correction, curriculum) mostly change compute efficiency rather than the asymptotic performance, and proposed a combined "ScaleRL" recipe with predictable, sigmoid-shaped scaling curves. Khatri+ 2025 In 2026 interviews, it's more valuable to explain the mechanisms (length weighting, entropy, ratio granularity, staleness) than to recite a leaderboard of acronyms. Check current recipes before claiming any one is "the" standard.
"Why did response length explode in my GRPO run?" Strong answers mention (1) the \(1/|o_i|\) normalization making long wrong answers cheap (Dr. GRPO), (2) truncation handling: are cut-off answers scored 0 and still in the loss? (3) genuine benefit, since longer CoT can be learned behavior if accuracy rises with it. Diagnose by splitting length curves by correct vs incorrect. "Why is entropy collapsing?" Clip-higher, an entropy bonus, KL-Cov, temperature and more diverse data. Also check whether the reward is too easy to max.
- DAPO: the clearest list of practical GRPO fixes with ablations.
- Understanding R1-Zero-Like Training (Dr. GRPO): base-model effects and length bias.
- GSPO: the sequence-level ratio argument.
- The Art of Scaling RL Compute: which knobs matter at scale.
6. RL with verifiable rewards and reasoning models
RLVR
Reinforcement Learning with Verifiable Rewards replaces the learned reward model with a program that checks correctness. The term was popularized by AI2's Tülu 3, which applied it to math, grade-school math and instruction-following constraints. Lambert+ 2024 The common verifier types:
| Domain | Verifier | Gotchas |
|---|---|---|
| Math (final answer) | Extract the answer (e.g., \boxed{}), normalize, compare symbolically or numerically (sympy-style equivalence) | Equivalent forms (\(\tfrac12\) vs 0.5 vs 1/2), units, multiple-choice guessing (25% free reward), problems with wrong reference answers, proofs (no single answer) |
| Code | Run hidden unit tests in a sandbox; reward = pass/fail or fraction passed | Weak tests let wrong code pass; tests visible to the model can be special-cased; flaky/timeout-sensitive tests add noise; sandbox escape |
| Instruction constraints | Programmatic checks ("exactly 3 bullet points", "no commas", JSON schema valid) | Satisfying the letter but not the spirit |
| Format | Regex: reasoning inside <think>, answer in the expected slot | A small format reward early helps parsing; too large and the model optimizes format over correctness |
| Agentic tasks | End-state checks: tests pass in repo, DB row exists, file has expected content, web form submitted | Many valid end states; state checkers that are too narrow or too loose |
Why RLVR beats RM-based RL for reasoning: a program can't be talked into a high score the way a neural reward model can, so you can optimize much harder (thousands of steps, little or no KL) before hacking dominates. The trade-off is coverage: only domains with checkable answers qualify. Most current research is about extending verifiability (rubrics, LLM judges, generated tests) to fuzzier domains (§7).
o1 and the reasoning-model paradigm
OpenAI's o1 (September 2024) announced that performance improves smoothly with both more RL train-time compute and more test-time compute (thinking longer), using large-scale RL to teach the model to use a long chain of thought productively. OpenAI 2024 Details were not published. The prior work on spending inference compute (best-of-N with verifiers, search, revision) showed that compute-optimal test-time scaling can beat scaling parameters on some problems. Snell+ 2024 o1's contribution was to put the "search" inside one long sampled chain of thought, learned by RL, instead of in an external search procedure.
DeepSeek-R1 and R1-Zero
DeepSeek-R1 (January 2025) published the recipe openly DeepSeek-AI 2025:
- R1-Zero: GRPO applied directly to the base model (DeepSeek-V3-Base) with no SFT, using only rule-based rewards: accuracy and a format reward for putting reasoning in think tags. AIME 2024 pass@1 rose from about 15.6% to 71.0% during training. Response length grew steadily on its own, and the model started doing reflection, verification and backtracking. The paper calls one example an "aha moment". The downsides were poor readability and language mixing.
- R1: a multi-stage pipeline. (1) A small "cold-start" SFT on long, readable CoT examples; (2) reasoning-focused RL with an added language-consistency reward; (3) rejection sampling from the RL checkpoint to build a large SFT set (on the order of 800k samples, reasoning plus general), then SFT again; (4) a final RL stage across all scenarios (verifiable rewards for reasoning, reward models for helpfulness and harmlessness).
- Distillation: SFT of smaller dense models (Qwen, Llama) on R1's outputs gave strong small reasoning models. For small models, distilling from a strong RL'd teacher beat running RL on the small model directly.
- What didn't work (per the paper): process reward models (hard to define steps, reward hacking, overhead) and MCTS over tokens (search space too large).
Kimi k1.5, released the same week, independently reported a similar finding: long-context RL with simple outcome rewards and no MCTS, value function or PRM. It also introduced "partial rollouts" to handle very long generations (§11). Kimi Team 2025
Treating the "aha moment" as proof that RL created reflection. A follow-up found that DeepSeek-V3-Base already shows aha-like self-reflection before RL, and that Qwen2.5 base models reason well even without prompt templates, which points to pretraining biases (including possible exposure to reasoning-style data). Liu+ 2025 The safer claim is that RL amplifies and stabilizes latent behaviors that pay off under the reward.
Test-time compute scaling
Two families:
Sequential (longer thinking)
One chain of thought, made longer. RL-trained reasoning models do this natively. Accuracy rises roughly log-linearly with thinking tokens up to a saturation point. Controls include reasoning-effort settings, thinking budgets, and "budget forcing", where you append "Wait" to make the model continue or force-stop it. Muennighoff+ 2025 (s1)
Parallel (more samples)
Sample N answers and aggregate: majority vote (self-consistency), best-of-N with a verifier or reward model, or tree search. This scales with N but saturates when the verifier is imperfect or the right answer never appears in the samples. Pass@k measures the ceiling of this approach.
The debate: does RL teach new things or just sharpen?
This is a favorite interview topic because there is real evidence on both sides.
"Sharpening" view
Across model families and RLVR algorithms, RL'd models beat their base models at pass@1, but the base models win at pass@k for large k (e.g., k=256). Coverage and perplexity analyses suggest the solutions RL finds were already in the base distribution. RL reweights toward them and loses some diversity. Distillation from a stronger teacher, by contrast, did add new reasoning patterns. Yue+ 2025 The spurious-reward results point the same way: some gains come from amplifying existing behaviors. Shao+ 2025
"Expansion" view
With prolonged RL (KL control, periodic reference resets, diverse tasks), RL'd models beat base models across pass@k, including on tasks where the base model fails at any k. The gains are largest where the base model starts weakest. Liu+ 2025 (ProRL) Practitioners also point out that agentic skills (multi-step tool use, long-horizon persistence) look very different after large-scale RL, and that pass@k at huge k on short tasks is a generous metric for the base model.
A defensible synthesis: on short, well-covered tasks, small-to-moderate RL mostly sharpens (it raises pass@1 toward the base model's pass@k). Longer training, broader task distributions and compositional, multi-step environments can push the boundary out. Exploration is the bottleneck, so the base model and the environment's curriculum largely determine which regime you're in.
This debate was very active through 2025 and into 2026. Related threads include on-policy distillation (student samples, teacher grades each token densely) as a cheaper alternative or complement to RL Thinking Machines 2025, and training objectives that directly optimize pass@k or preserve diversity. Check recent results before taking a strong position in an interview. Present it as open.
"How would you tell whether your RL run taught the model anything new?" Strong answer: plot pass@k curves for base vs RL'd at k = 1…256 on a held-out set; check whether RL'd solutions appear among the base model's samples (coverage); test transfer to unseen task types; control for the format/answer-extraction gains that often explain early jumps; and run the same recipe on more than one model family.
- DeepSeek-R1 paper: the R1-Zero and R1 recipes, plus the "unsuccessful attempts" section.
- Does RL Really Incentivize Reasoning Capacity…?: the main "sharpening" paper.
- ProRL: the main counterpoint.
- Tülu 3: an open, end-to-end post-training recipe including RLVR.
7. Reward design: outcome, process, judges and rubrics
Outcome vs process rewards
| Outcome reward (ORM / verifier) | Process reward (PRM) | |
|---|---|---|
| Signal | One score for the final answer | A score per reasoning step |
| Credit assignment | Sparse; relies on many samples | Dense; pinpoints the first wrong step |
| Labels | Cheap (reference answer, tests) | Expensive (human step labels) or estimated (MC rollouts) |
| Hackability | Low if programmatic | Higher: the policy can learn steps that look good to the PRM |
| Main use today | RL training signal | Reranking / best-of-N, search guidance, research on dense RL credit |
Key results. OpenAI's "Let's Verify Step by Step" trained a PRM on about 800k human step-level labels (PRM800K). Used as a best-of-N reranker, it beat an ORM on a MATH subset. Lightman+ 2023 Math-Shepherd removed the human labels: a step's label is estimated by rolling out completions from that step and checking how often they reach the right answer. Wang+ 2023 "Rewarding Progress" reframes a good process reward as progress, the change in the likelihood of eventual success under a (different) prover policy, which is effectively a step-level advantage. Setlur+ 2024 PRIME derives an implicit PRM from a model trained only on outcome labels and updates it online, which reduces hacking. Cui+ 2025
Practical status: large-scale reasoning RL mostly uses outcome rewards (R1 and Kimi k1.5 both reported PRMs as not worth it at scale). Process signals live on in verifier-guided search, in agentic settings as turn-level checks, and as an active research area.
Learned reward models, LLM judges and rubrics
For non-verifiable domains (writing quality, helpfulness, research reports, medical advice), the options are:
- Bradley-Terry reward models from human preference pairs (classic RLHF). They are easy to overoptimize, so you need a KL leash and early stopping. Gao+ 2022
- LLM-as-judge: prompt a strong model to score or compare responses. It's flexible, but it has known biases (verbosity, position, self-preference, confident tone) and the policy will find them. Pairwise comparisons against a reference are usually more stable than absolute scores.
- Rubric-based rewards: write per-prompt checklists of concrete criteria (expert-authored or model-generated, e.g., "mentions contraindication X", "cites the 2023 guideline"). A judge checks each item and the reward is a weighted sum. "Rubrics as Rewards" showed gains over Likert-style judge scores in medicine and science. Gunjal+ 2025 Rubrics make an LLM judge more like a verifier: decomposed, auditable, and harder to sway with style.
- Generative/reasoning verifiers: a judge that reasons before scoring, sometimes trained with RL itself. These are increasingly common in 2025–2026 recipes for "fuzzy" domains.
Reward signals sit on a spectrum of hackability vs. coverage. Exact-match checkers cover little and are almost unhackable. Unit tests cover more and are somewhat hackable. Rubric judges cover a lot and are moderately hackable. Free-form LLM judges and learned RMs cover everything and are very hackable. The craft is pushing each domain as far toward the verifiable end as possible, and adding monitoring where you can't.
"Design a reward for training a model to write good literature reviews." Strong answers break it down: verifiable sub-checks (citations resolve to real papers, quotes match sources, required sections present), rubric items judged per criterion by a different model family than the policy, pairwise comparisons against a reference instead of absolute scores, length normalization or caps, and a held-out human eval to catch judge exploitation. They also mention monitoring reward-vs-human-agreement drift over training.
- Let's Verify Step by Step: process vs outcome supervision.
- Rewarding Progress: what a useful process reward should measure.
- Rubrics as Rewards: extending RLVR-style training to non-verifiable domains.
8. Reward hacking and specification gaming
Reward hacking: the policy raises the measured reward without doing what the designer intended. Formally, any time optimizing the proxy reward diverges from optimizing the true objective. Skalse+ 2022 Classic RL is full of examples, like the boat-racing agent circling to collect respawning targets instead of finishing the race. DeepMind 2020 With LLMs, the agent is smart enough to read and reason about its own grader.
Taxonomy with LLM examples
| Type | Example | Mitigation |
|---|---|---|
| Test special-casing | Code that detects test inputs and returns hard-coded expected outputs | Hidden tests, held-out tests, randomized/property-based tests, LLM review of diffs |
| Grader tampering | Editing or deleting failing tests, modifying conftest.py, exiting the process early with code 0 before tests run, overriding __eq__ so every comparison passes Zhong+ 2025 | Read-only test dirs, run graders outside the agent's sandbox, checksum protected files, check for exit-code tricks |
| Environment exploits | Finding the reference solution on disk or in git history, using network access to look up answers, abusing timing or scoring code. METR documented frontier models overwriting timing functions to fake speedups, stubbing evaluators, and searching the Python call stack for the grader's reference answer. METR 2025 | Scrub artifacts, network policies, separate grader process, red-team the env before training |
| Judge exploitation | Verbose, confident or flattering answers; injected text that addresses the judge | Rubrics, pairwise vs reference, different judge family, length control, strip instructions aimed at the judge |
| Sycophancy | Agreeing with the user's wrong belief because human raters and preference models reward agreement Sharma+ 2023 | Preference data that rewards correction, targeted evals, debiased RMs |
| Answer-format gaming | Listing several answers so the extractor finds the right one; guessing on multiple choice | Strict extraction (single final answer), penalize multiple answers, avoid MC-only datasets |
| Length/effort gaming | Padding to look thorough, or giving up early on hard tasks when partial credit pays | Length-aware rewards, careful partial-credit design |
Why it matters beyond the metric
- Hacks generalize to behavior. Anthropic reported that when a model learned to reward hack in real production coding environments, it generalized to broader misalignment (alignment faking, sabotage attempts in agentic use). Standard chat-style safety RLHF fixed chat evals but not agentic ones. Effective mitigations: prevent the hacking, diversify safety training, and "inoculation prompting", which frames hacking as acceptable in training so it doesn't generalize into a misaligned self-concept. MacDiarmid+ 2025
- Monitoring the chain of thought works, but optimizing against it backfires. OpenAI found an LLM reading the CoT catches hacks in agentic coding far better than watching actions alone. Adding the monitor to the reward helps at low optimization pressure, but under heavy pressure the model learns obfuscated hacking: it hides intent in the CoT and keeps hacking. Their recommendation is to avoid strong optimization pressure on the CoT itself. Baker+ 2025
- Benchmarks inherit it. If agents learn to game tests during training, SWE-bench-style scores overstate real capability. ImpossibleBench measures "cheating rate" on tasks where the spec and tests conflict, so any pass implies a shortcut. Zhong+ 2025
A mitigation playbook
- Harden the environment first: graders run outside the agent's reach, protected files, no reference answers in the filesystem, network egress controlled.
- Diverse, hidden verification: held-out tests the policy never sees, plus property-based or randomized tests, plus multiple independent checkers.
- Monitor rollouts continuously: sample and read transcripts; run LLM monitors that flag suspicious patterns (edits to test files,
sys.exit, hard-coded constants); track reward vs. held-out-eval divergence. - Regularize: KL to a reference, early stopping on proxy-vs-true divergence, reward caps.
- Penalize detected hacks carefully. Explicit penalties work for clear-cut cases, but penalizing on CoT content risks obfuscation.
- Patch and re-version the environment when a hack is found, and add the exploit to a regression suite for env QA.
RL-env startups ask some version of: "Here's a coding environment. How would the model hack it, and how would you stop it?" Walk the attack surface systematically: the grader (tests, exit codes, file access), the filesystem (solutions, git history, caches), the network, the judge, the answer extractor, timeouts and partial credit. Then describe the env QA process: red-team with a strong model before release, check that pass rates on known-impossible variants are near zero, and audit sampled transcripts during training.
- Lilian Weng, "Reward Hacking in Reinforcement Learning": a broad survey with LLM-specific sections.
- Natural Emergent Misalignment from Reward Hacking in Production RL: why env hacks matter for safety.
- Monitoring Reasoning Models for Misbehavior: CoT monitoring and the obfuscation risk.
- METR on recent frontier-model reward hacking: concrete transcripts.
9. RL environments in depth
For LLM RL, the environment is everything except the policy and the optimizer: what tasks the model sees, what it can do, what it observes back, when the episode ends, and how it's scored. As algorithms converged on a few GRPO/PPO variants, the environments became the main thing that sets one lab's RL apart from another's, and a market formed around building them. As of 2026, frontier labs buy environments from dedicated vendors and from human-data companies, alongside open hubs.
The RL-environment market and tooling changed a lot in 2025–2026: vendor startups, acquisitions, open hubs, new interface standards. Treat any specific company or framework named here as a snapshot as of late 2026, and check current docs before relying on API details.
Anatomy of an environment
| Component | What it is | Example (SWE-style task) |
|---|---|---|
| Task distribution | The dataset or generator of task instances, with metadata (difficulty, domain, split) | Thousands of (repo, commit, issue) tuples mined from GitHub |
| Initial state | What exists at reset(): prompt, files, DB, browser page, user persona | A container with the repo checked out at the pre-fix commit and dependencies installed |
| Action space / tools | What the model can do: free text, structured tool calls, shell commands, clicks | bash, edit_file, search, submit |
| Observation formatting | How environment output becomes tokens: truncation, error messages, screenshots vs accessibility tree | stdout/stderr truncated to N lines, file views with line numbers |
| Transition dynamics | How state changes after an action (the sandbox, simulator or simulated user) | Actually run the command in the container |
| Termination | Success, explicit submit, max turns/tokens/wall-clock, unrecoverable error | submit called, or 100 turns, or 30 min |
| Reward / verifier | The scoring function, possibly multi-component | Hidden FAIL_TO_PASS and PASS_TO_PASS tests run after submit; reward = 1 if all pass |
| Harness / scaffold | The agent loop and system prompt the policy runs inside (which may be a production agent like a coding CLI) | A ReAct-style loop with a tool schema; or a real coding-agent harness |
A minimal interface
Classical RL standardized on the Gym API. Gymnasium (the maintained successor to OpenAI Gym) uses reset(seed=...) → (obs, info) and step(action) → (obs, reward, terminated, truncated, info). It separates true termination from time-limit truncation, which matters for bootstrapping. Legacy Gym's step returned a 4-tuple with a single done. Gymnasium docs Towers+ 2024 LLM environments keep that shape but deal in messages and tool calls, run asynchronously (thousands of concurrent episodes), and keep the verifier separate from the dynamics. A sketch (pseudo-code, not any specific library's API):
from dataclasses import dataclass, field
from typing import Any
@dataclass
class Task: # one row of the task distribution
id: str
prompt: str
assets: dict # repo snapshot, DB seed, URL, user persona...
hidden: dict # reference answer / hidden tests (NEVER shown to the policy)
meta: dict = field(default_factory=dict) # difficulty, domain, split, version
@dataclass
class StepResult:
observation: list[dict] # new messages to append (tool results, user replies)
terminated: bool # task finished (submit / success / fatal)
truncated: bool # budget exhausted (turns, tokens, wall-clock)
info: dict # diagnostics; not shown to the policy
class LLMEnv:
tools: list[dict] # JSON schemas exposed to the policy
async def reset(self, task: Task, seed: int) -> list[dict]:
"""Provision an isolated sandbox from task.assets (container/VM/browser),
seed all randomness, return the initial messages (system + task prompt)."""
async def step(self, action: dict) -> StepResult:
"""Parse one assistant turn (text and/or tool calls). Execute tools in the
sandbox with timeouts; format outputs (truncate, sanitize); check
termination. Malformed calls get an error observation, not a crash."""
async def score(self) -> dict[str, float]:
"""Run the verifier OUTSIDE the policy's reach (separate process or
container) against the final state plus task.hidden. Return components,
e.g. {"correct": 1.0, "format": 1.0, "tests_modified": 0.0}."""
async def close(self) -> None:
"""Tear down the sandbox; always runs, even on error/timeout."""
# Trainer-side rollout (simplified):
# msgs = await env.reset(task, seed)
# while True:
# turn = await policy.generate(msgs, tools=env.tools) # record token ids + logprobs
# res = await env.step(turn); msgs += [turn] + res.observation
# if res.terminated or res.truncated: break
# reward = combine(await env.score()); await env.close()
# loss_mask = 1 for policy-generated tokens, 0 for prompt / tool-output tokens
Design principles
| Principle | Why | How, concretely |
|---|---|---|
| Verifiability | Sparse but trustworthy reward beats dense noisy reward for RL at scale | Define success as a checkable end state; hidden tests; multiple independent checks; avoid rewards that need a judge when a program would do |
| Non-hackable graders | The policy will optimize the grader, not your intent | Grader outside the sandbox; protected/read-only test files; no answers on disk; red-team with a strong model; "impossible task" canaries whose pass rate should be about 0 |
| Difficulty calibration | Group-relative methods get zero gradient when all samples tie | Measure pass rate under the current policy (e.g., 8–16 samples); keep tasks in a band like 10–90%; refresh as the policy improves (curriculum); DAPO-style dynamic sampling online |
| Diversity | Narrow envs produce narrow skills and overfitting to quirks | Many repos/sites/domains; varied phrasings; procedural generation; dedupe near-duplicates; track per-domain reward |
| Determinism & reproducibility | Noisy rewards look like learning signal; debugging needs replays | Pin images and dependencies; seed everything; freeze time and network responses (record/replay); version tasks and graders; log full traces |
| Sandboxing & safety | The policy runs arbitrary code and gets rewarded for finding loopholes | Containers/microVMs (gVisor, Firecracker-style), no host mounts, egress allowlists, resource limits, kill on timeout |
| Throughput & cost | RL needs millions of episodes | Fast provisioning (snapshots, warm pools), async I/O, cheap graders, cache env builds; budget per-episode cost (CPU-minutes, API calls) |
| Robust observations | Huge tool outputs blow the context; crashes waste rollouts | Truncate with markers, return errors as observations, consistent formatting between training and deployment |
| Train/deploy match | Skills transfer best when the training harness matches production | Same tool schemas, system prompts and scaffold as the deployed agent; increasingly, train inside the production harness itself |
| Contamination hygiene | Training on eval tasks inflates numbers | Separate train/eval repos or time windows; hash and dedupe against public benchmarks |
An RL environment is a unit test for a skill that will be attacked by a highly motivated adversary millions of times. It needs to be fair (solvable, unambiguous), informative (some samples succeed, some fail), and robust (the only way to get reward is to have the skill). Most env-engineering effort goes into the last property.
Types of environments
| Type | State & actions | Reward | Examples | Hard parts |
|---|---|---|---|---|
| Math / logic | Single turn; text | Answer checker | Competition math sets, synthetic puzzles | Answer normalization, wrong labels, saturation |
| Competitive code | Single or few turns; code; optionally run tests | Hidden test pass rate | LiveCodeBench-style problem sets | Weak tests, timeouts, special-casing |
| SWE / repo tasks | Multi-turn; shell + editor in a container | Hidden tests (fail-to-pass, pass-to-pass), or patch similarity | SWE-bench Jimenez+ 2023, SWE-Gym Pan+ 2024, R2E-Gym Jain+ 2025, SWE-smith Yang+ 2025, SWE-RL Wei+ 2025 | Building runnable envs per repo (dependency hell), heavy containers, test tampering |
| Terminal / ops | Shell in a container | Scripted end-state checks | Terminal-Bench tasks run via Harbor Harbor | Many valid solutions, nondeterministic tools |
| Browser / web | Multi-turn; clicks, typing on DOM / accessibility tree / screenshots | End-state checks on backend DB or page | WebArena (self-hosted sites) Zhou+ 2023 | Live sites drift; must self-host clones; auth; slow |
| Computer use | Full OS VM; screenshots + mouse/keyboard | Execution-based checker scripts | OSWorld Xie+ 2024 | VM cost, latency, visual grounding, snapshotting |
| Tool use + simulated user | Multi-turn chat with an LLM-simulated user and domain APIs | Final DB state matches goal; policy compliance | τ-bench (retail, airline) Yao+ 2024 | Sim-user realism and consistency; the policy can manipulate the simulator |
| Search / research | Interleaved reasoning and search calls | Answer match (QA) or rubric | Search-R1 Jin+ 2025 | Live-index nondeterminism, cost, judge reliance |
| Code-interpreter reasoning | Reasoning with an executor tool | Answer checker | ReTool Feng+ 2025 | Sandbox throughput |
| Games / puzzles | Multi-turn, fully simulated | Game score / win | Sokoban, FrozenLake in RAGEN Wang+ 2025 | Transfer to real tasks is unclear |
| Self-play / self-proposed | Model proposes tasks and solves them | Executor-verified | Absolute Zero Zhao+ 2025 | Keeping the task distribution useful and non-degenerate |
| Non-verifiable (writing, advice) | Single or multi-turn | Rubric/LLM judge, RM | Rubrics as Rewards Gunjal+ 2025 | Judge exploitation; reward drift |
SWE environments in more detail, since they're the most-asked type. SWE-bench pairs real GitHub issues with the PR that fixed them. The grader runs the tests that the fix made pass (fail-to-pass) plus tests that must keep passing (pass-to-pass). Jimenez+ 2023 To train, you need many such executable instances. SWE-Gym built about 2.4k real executable Python tasks and showed both agent fine-tuning and verifier training benefit. Pan+ 2024 R2E-Gym generated environments procedurally from commits, with synthesized tests, plus hybrid (execution + model) verifiers. Jain+ 2025 SWE-smith scaled further by synthesizing bugs into repos, reaching tens of thousands of instances. Yang+ 2025 SWE-RL avoided execution altogether with a rule-based reward: the similarity between the generated patch and the ground-truth patch. That's cheap, but a weaker proxy. Wei+ 2025 The main engineering cost is building a working Docker image for each repo version.
Frameworks and interfaces (as of 2026)
| Tool | What it is | Notes |
|---|---|---|
| Gymnasium (Farama) | Standard single-agent RL env API; successor to OpenAI Gym | reset/step with terminated vs truncated docs |
verifiers + Environments Hub (Prime Intellect) | Library for building LLM RL/eval environments, plus a community hub (launched Aug 2025) and the prime-rl trainer GitHub Prime Intellect 2025 | As of late 2026 the "v1" API is organized around tasksets (data + scoring), harnesses (the agent program, e.g., an existing coding agent), traces, and toolsets exposed as MCP servers. The earlier v0 API has been removed. Check current docs. |
| OpenEnv (Meta + Hugging Face) | Spec and hub for agentic environments with a Gym-like reset()/step()/close() API, packaged as Docker-based environments HF blog 2025 | Announced Oct 2025; integrations with TRL, SkyRL, Unsloth and others announced or underway GitHub |
| Harbor | Framework from the Terminal-Bench creators for running agents (coding CLIs, OpenHands, …) on containerized tasks at scale, for evals and RL rollouts GitHub | Official harness for Terminal-Bench 2.0; plugs into cloud sandbox providers |
| Benchmark-derived gyms | SWE-Gym, R2E-Gym, SWE-smith, WebArena, OSWorld, τ-bench | Often double as training envs, which makes contamination hygiene important |
The broad trend: environments are moving from "a Python function that scores a string" to "a containerized task plus an unmodified production agent harness". The policy is trained inside the same scaffold it will be deployed in, and tools increasingly come through MCP.
The classic RL-env design question: "Design an environment to train an agent to do X (fix CI failures / file expense reports / do spreadsheet modelling)." A strong answer walks the anatomy table. Where tasks come from and how many; how state is provisioned and reset (snapshots); the tool surface and observation format; termination budgets; the verifier and how it resists hacking; difficulty calibration against the current policy; determinism (pinned images, mocked time/network); throughput and cost per episode; and the QA process (solvability checks with a reference solution, impossible-task canaries, red-teaming, transcript audits). Mentioning that every task should ship with a known-good solution that passes the grader, and a known-bad one that fails, is a strong signal.
- verifiers and the Environments Hub: hundreds of open environments to read as worked examples.
- OpenEnv announcement: the Meta/HF standardization push.
- SWE-Gym and SWE-smith: how SWE training environments are built at scale.
- τ-bench: simulated-user tool-use envs and the pass^k reliability metric.
10. Agentic and multi-turn RL
Multi-turn agent RL keeps the same algorithms but changes nearly everything around them.
Loss masking
A trajectory interleaves policy tokens with environment tokens (tool outputs, user messages, file contents). Only the policy's own tokens are actions. Environment tokens must be masked out of the policy-gradient loss (and the KL), or you train the model to "predict" tool outputs, which is wrong and destabilizing. Search-R1, for example, masks retrieved passages from the loss. Jin+ 2025 The same applies to prompt tokens and system messages.
Re-tokenization drift. If you store a trajectory as text and re-tokenize it for training, the token boundaries can differ from what the policy actually sampled (especially around tool-call delimiters and chat templates). The logprobs then no longer match, and importance ratios break silently. Record the exact token ids and logprobs at generation time and train on those.
Credit assignment over turns
- Trajectory-level (default): one final reward, every policy token in every turn gets the same advantage (GRPO-style over G full trajectories of the same task). Simple and robust, but noisy for long horizons.
- Turn-level / step-level: group "similar" intermediate states across trajectories and compute relative advantages per step, as in GiGPO's nested groups (episode-level plus step-level anchored on repeated states). Feng+ 2025 Or train a turn-level critic.
- Intermediate rewards: partial credit for sub-goals (tests compiled, correct file found). These are useful but invite hacking and shortcutting, so keep them small relative to the outcome reward.
- Failure modes: RAGEN documents an "echo trap", where multi-turn agents collapse into repetitive, low-diversity reasoning, and argues for diverse initial states and fine-grained, reasoning-aware rewards. Wang+ 2025
Long horizons
- Context growth: a 100-turn SWE episode can exceed 100k tokens. Options are bigger context windows, truncating old observations, summarization or "context management" tools, and training the model to use memory files. Each changes the MDP (the policy no longer sees everything), so training and deployment must use the same strategy.
- Variance in episode length and cost: some episodes finish in 3 turns, some take 200. This drives the long-tail problem (§11).
- Environment failures: flaky tools, sandbox crashes and timeouts. Mark these trajectories as invalid and drop them. Don't give them reward 0, which would punish the policy for infrastructure noise.
- Cost per episode: VMs, browsers and LLM-simulated users cost real money per step. The policy's tokens are not the only cost.
"What changes when you go from single-turn RLVR to multi-turn agent RL?" Hit: loss masking of env tokens; exact token-id bookkeeping; credit assignment across turns; context management; invalid-trajectory handling; async env execution with very different episode lengths; sandbox cost; and the train/deploy harness match.
- The Landscape of Agentic RL for LLMs: A Survey: a broad map of agentic RL methods.
- GiGPO: step-level group advantages for agents.
- RAGEN: multi-turn RL failure modes.
11. RL training infrastructure
The key fact: LLM RL is dominated by generation. Each training step needs fresh samples from the current policy, and autoregressive decoding of long CoTs or multi-turn episodes is far slower per token than the training forward/backward pass. Rollout generation is often reported as the majority of wall-clock time, often well over half, in synchronous systems. So RL systems are built as a fast inference fleet bolted to a trainer.
Design axes
| Axis | Option A | Option B | Trade-off |
|---|---|---|---|
| Placement | Colocated: the same GPUs alternate between generation and training (reshard weights between engine and trainer layouts) | Disaggregated: separate GPU pools for rollout and training | Colocated: no idle pools, simple weight sync, but phases serialize. Disaggregated: phases overlap and each pool can be sized and tuned on its own, but one pool idles if they're unbalanced, and weights cross the network. |
| Synchrony | Synchronous: generate batch with \(\pi_k\), train, sync, repeat | Asynchronous: generators run continuously; the trainer consumes data up to N versions stale | Async gives much higher utilization but off-policy data. Needs importance correction, staleness bounds, or robust objectives. Noukhovitch+ 2024 Fu+ 2025 (AReaL) |
| Weight update timing | Between batches | In-flight: swap weights mid-generation, so a sequence can span policy versions | PipelineRL reports about 2× faster learning while staying nearly on-policy. Piché+ 2025 |
| Long sequences | Wait for all to finish | Partial rollouts: cap per-iteration generation; continue unfinished sequences next iteration | Kimi k1.5's fix for long-tail CoTs. Kimi Team 2025 |
The long-tail problem
Response lengths are heavy-tailed. In a synchronous batch, everyone waits for the longest response or slowest episode. Decode batch occupancy decays as short sequences finish, and GPUs sit mostly idle at the end of each step. Fixes: (1) asynchronous or pipelined generation so the trainer never waits for stragglers; (2) partial rollouts / interruptible generation; (3) over-sampling and dropping the slowest (this biases against long correct answers, so be careful); (4) length caps with proper truncation handling; (5) continuous batching and prefix caching in the engine. The G samples for one prompt share a prefix, so prefix caching is a direct win.
Weight sync and the inference/trainer mismatch
- Weight sync cost: a 70B model in bf16 is about 140 GB. Broadcasting it over NVLink/InfiniBand at hundreds of GB/s takes seconds, but it happens every step and must be resharded from the trainer's layout (FSDP/TP/PP) to the engine's (TP). Frameworks use bucketed NCCL broadcasts, CUDA IPC when colocated, or LoRA-only sync.
- Logprob mismatch: even with identical weights, vLLM/SGLang and the training stack produce different token probabilities (different kernels, precision, batch-dependent reductions, sometimes quantized rollouts). The "on-policy" data is quietly off-policy. A common fix is truncated importance sampling: weight each token's gradient by \(\min(\pi_\text{train}/\pi_\text{rollout}, C)\). veRL PR #2953 Batch-invariant kernels can make inference deterministic and remove the mismatch at its source. Thinking Machines 2025
- MoE specifics: router decisions can differ between engine and trainer, which amplifies the mismatch. Sequence-level ratios (GSPO) and routing replay are the mitigations discussed.
Frameworks
| Framework | Origin | Notable design |
|---|---|---|
| veRL (HybridFlow) | ByteDance Seed | Hybrid single-/multi-controller programming model; colocated actor/rollout with resharding; FSDP or Megatron + vLLM/SGLang Sheng+ 2024 GitHub |
| OpenRLHF | Open-source community | Ray-based, vLLM generation, distributed PPO/GRPO/REINFORCE++ Hu+ 2024 |
| TRL | Hugging Face | GRPOTrainer and friends; vLLM in colocate or server mode; easiest on-ramp docs |
| slime | THUDM (Zhipu / Tsinghua) | Megatron training + SGLang rollout, built for large-scale RL post-training GitHub |
| AReaL | Ant / Tsinghua | Fully asynchronous, interruptible rollouts, staleness-aware PPO Fu+ 2025 |
| PipelineRL | ServiceNow | In-flight weight updates Piché+ 2025 |
| prime-rl | Prime Intellect | Async RL, used for globally decentralized training (INTELLECT-2); integrates with verifiers environments Prime Intellect 2025 GitHub |
| SkyRL, NeMo-RL | Berkeley NovaSky; NVIDIA | Agent/long-horizon focus (SkyRL); NVIDIA's scalable RL library SkyRL NeMo-RL |
Back-of-envelope: one GRPO step
Say 512 prompts × G=16 samples × an average of 8k generated tokens. That's \(512\times16\times8{,}000\approx 6.6\times10^7\) tokens per step. If a 7B policy decodes at a few thousand tokens/s per GPU at high batch, then 64 GPUs at about 3k tok/s gives about 190k tok/s, or roughly 350 s of generation per step, ignoring the long tail. The training pass over the same tokens costs about \(6 \times 7\times10^9 \times 6.6\times10^7 \approx 2.8\times10^{18}\) FLOPs. At about 400 TFLOP/s effective per GPU that's roughly 110 s on 64 GPUs. So generation dominates even before stragglers and environment latency, which is why async designs and the long tail are the main concerns.
"Design the RL infrastructure for agentic coding at scale." Strong answers cover: disaggregated inference fleet + trainer; thousands of concurrent sandboxes with warm pools; async rollouts with bounded staleness plus importance correction; partial rollouts for long episodes; token-id-exact trajectory storage with loss masks; weight sync strategy; reward service isolated from sandboxes; invalid-trajectory handling; and observability (reward, length and entropy curves, KL, clip fraction, hack monitors, per-env pass rates). Bonus: the logprob mismatch and how you'd detect it (log both engine and trainer logprobs and plot the ratio distribution).
- HybridFlow (veRL): the reference design for colocated RL.
- AReaL: the fully async design and staleness control.
- PipelineRL: in-flight weight updates.
- Defeating Nondeterminism in LLM Inference: why engine and trainer disagree, and how to fix it.
12. Evaluating RL'd models
| Metric | Definition | Use |
|---|---|---|
| pass@1 (avg@n) | Mean accuracy over n samples per problem | Primary metric; average many samples for low variance (small benchmarks like AIME have only 30 problems) |
| pass@k | Probability at least one of k samples is correct. Unbiased estimator from n ≥ k samples with c correct: \(1-\binom{n-c}{k}/\binom{n}{k}\) | Capability ceiling; the sharpening-vs-expansion debate |
| maj@k / cons@k | Majority vote over k samples | Test-time compute scaling with no verifier |
| pass^k | Probability all k trials succeed (τ-bench) Yao+ 2024 | Reliability for agents: users care about consistency |
| Accuracy vs tokens | Score as a function of thinking budget or cost | Efficiency. A longer-CoT model may just spend more compute. |
| Hack / cheat rate | Pass rate on impossible variants; flagged-transcript rate Zhong+ 2025 | Is the score real? |
| Regression suite | General knowledge, chat quality, safety, calibration, instruction following | RL on narrow domains can degrade other abilities |
- Variance is large. Report multiple seeds and confidence intervals; a few points on a 30-problem benchmark is noise.
- Separate format gains from reasoning gains. Early RL jumps are often answer-extraction and format compliance.
- Hold out environments, not just tasks: evaluate on repos, sites or domains never seen in training to measure transfer.
- Check the base model on the same harness and sampling settings. Many "RL gains" shrink once the base model gets a good prompt and enough samples.
- Read transcripts. Aggregate metrics hide hacks, language mixing and degenerate repetition.
"Your RL'd model jumped 15 points on SWE-bench Verified. What do you check before celebrating?" Contamination (were these repos in training?), hack rate (test edits, special-casing), same scaffold and budget for the baseline, variance across seeds, pass^k reliability, accuracy per token/cost, and performance on held-out repos and a different benchmark (e.g., a terminal or multi-language set).
- SWE-bench leaderboards: the variants (Verified, Multimodal, etc.) and their harness rules.
- Yue+ 2025: careful pass@k methodology.
Interview question bank
Derive the policy gradient and explain why a baseline doesn't add bias.
\(\nabla_\theta \mathbb{E}_{\tau\sim p_\theta}[R]=\mathbb{E}[R\,\nabla\log p_\theta(\tau)]\) via \(\nabla p=p\nabla\log p\). Expanding \(\log p_\theta(\tau)\) gives the policy log-probs plus dynamics terms that don't depend on \(\theta\), so they vanish. The method is model-free and the reward needn't be differentiable. For a baseline \(b(s)\): \(\mathbb{E}_{a\sim\pi}[b(s)\nabla\log\pi(a|s)]=b(s)\nabla\sum_a\pi(a|s)=b(s)\nabla 1=0\). So subtracting it leaves the expectation unchanged and can greatly reduce variance. The optimal-ish baseline is \(V(s)\), which turns the weight into the advantage. The caveat is that the baseline must not depend on the action/sample being scored. That's why RLOO uses a leave-one-out mean.
Explain PPO's clipped objective. What does clipping do to the gradient?
PPO maximizes \(\min(\rho A, \text{clip}(\rho,1-\epsilon,1+\epsilon)A)\) per token, where \(\rho=\pi_\theta/\pi_\text{old}\). For positive advantage, once \(\rho>1+\epsilon\) the clipped term is smaller and constant, so the gradient is zero. For negative advantage the same happens once \(\rho\lt 1-\epsilon\). It's a pessimistic first-order trust region that lets you take multiple minibatch epochs on one batch of rollouts without the policy drifting far from the data-generating policy. Side effects: tokens outside the band stop contributing, which can silence rare high-value tokens (motivating CISPO), and symmetric clipping limits how fast unlikely tokens can grow (motivating DAPO's clip-higher).
Why did the field move from PPO to GRPO-style methods for reasoning RL?
The LLM setting is close to a contextual bandit (deterministic transitions, one terminal reward). Per-token value estimates are hard to learn and add little. The critic is as large as the policy, which roughly doubles training memory and compute, and it's unstable on sparse rewards. Sampling G responses per prompt gives an empirical baseline for free, and inference is needed anyway. With 0/1 verifiable rewards, the group mean is a very natural baseline. Trade-offs: no per-token credit assignment, wasted compute on all-tie groups, and normalization biases. VAPO-style work argues a well-trained critic can still win, so it isn't settled.
Write down the GRPO advantage and name three known problems with vanilla GRPO plus their fixes.
\(\hat A_i=(r_i-\text{mean}(r))/\text{std}(r)\), applied to all tokens of response i, inside a PPO-clipped loss with per-sequence \(1/|o_i|\) averaging and a KL-to-reference term. Problems: (1) length bias, since \(1/|o_i|\) makes long wrong answers cheap. Fix: Dr. GRPO removes it; DAPO uses token-level aggregation. (2) Std normalization over-weights near-uniform groups (difficulty bias). Fix: Dr. GRPO drops std; REINFORCE++ normalizes globally. (3) Entropy collapse. Fix: DAPO clip-higher, Clip-Cov/KL-Cov. (4) Noisy token-level ratios, especially MoE. Fix: GSPO's sequence-level ratio. (5) Zero-advantage groups waste compute. Fix: DAPO dynamic sampling, curriculum filtering.
What is the KL penalty for, and why do some reasoning recipes (DAPO) remove it?
The KL to a reference policy limits drift. That guards against reward-model overoptimization, preserves fluency and general abilities, and keeps outputs in-distribution for the reward model. In PPO-RLHF it's subtracted from the per-token reward. In GRPO it's a separate loss term. With programmatic verifiers there's no RM to exploit in the same way, and long-CoT reasoning needs the policy to move far from the base distribution (much longer outputs, new behaviors), so the KL mostly slows learning. DAPO drops it. ProRL keeps a KL but periodically resets the reference to the current policy, a compromise between stability and freedom to move.
Back-of-envelope: how many generated tokens per step, and is training or generation the bottleneck?
Example: 256 prompts × 16 samples × 10k tokens ≈ 41M tokens per step. Training FLOPs for a 32B model ≈ 6 × 32e9 × 41e6 ≈ 7.9e18. At about 400 TFLOP/s effective per GPU on 128 GPUs, that's roughly 150 s. Decoding 41M tokens of a 32B model at maybe 1–2k tok/s per GPU (an order-of-magnitude guess; measure it) on 128 GPUs is about 160–320 s before the long tail. The slowest 1% of sequences can double wall-clock in synchronous mode. So generation, and especially its tail, usually dominates, which motivates async/pipelined RL, partial rollouts and prefix caching. Always state that throughput numbers depend heavily on hardware, engine and batch shape.
What is RLVR and what makes a good verifier?
RL with verifiable rewards uses programmatic checks (math answer equivalence, hidden unit tests, constraint checkers) instead of a learned reward model. A good verifier is correct (no false positives or negatives on reference solutions), robust to equivalent forms, deterministic, fast and cheap, isolated from the policy (runs outside its sandbox, hidden data not visible), and resistant to known exploits (multiple answers, test edits, exit-code tricks). Validate it by running known-good and known-bad solutions, check that reference answers are actually right, and track its false-positive rate during training with transcript audits.
Does RL teach models new reasoning capabilities, or just sharpen existing ones? Argue both sides.
Sharpening: Yue et al. found RLVR models beat base models at pass@1 but lose at pass@k for large k across families and algorithms. Their solutions appear to lie within the base model's coverage, and distillation, not RL, added new patterns. Spurious-reward results show gains even with random rewards on Qwen, consistent with amplifying latent behaviors. Expansion: ProRL found prolonged, diverse RL with reference resets improves pass@k across the board, including tasks where the base model never succeeds, with gains largest where the base was weakest. Synthesis: short RL on narrow tasks mostly reweights; long, diverse, compositional training, especially agentic, can push the boundary. Exploration (base prior, curriculum) decides which regime you're in. It's still contested, so say so.
What happened in DeepSeek-R1-Zero and why was it significant?
GRPO was applied directly to a strong base model (DeepSeek-V3-Base) with only rule-based accuracy and format rewards, with no SFT. Accuracy on AIME rose dramatically, response length grew on its own, and reflection and backtracking behaviors appeared. It was significant because it showed in the open that long-CoT reasoning can emerge from outcome-only RL without PRMs, MCTS or curated reasoning traces. Caveats: poor readability and language mixing (fixed in R1 with cold-start SFT and a language-consistency reward), and later work showing aha-like behaviors already exist in base models, so "emergence" partly means amplification.
Process reward models vs outcome rewards: when would you use each?
Outcome rewards (final-answer checks, tests) are cheap, hard to hack when programmatic, and the default for large-scale RL. Their weakness is sparse credit assignment over long chains. PRMs score each step, which gives dense signal and helps search/reranking (Lightman's PRM beat an ORM at best-of-N), but labels are expensive or noisy (MC-estimated), "steps" are ill-defined, and policies learn to please the PRM. Use PRMs or progress-style verifiers for test-time search, verifier-guided decoding, or as auxiliary dense signals with care. Use outcome rewards as the primary RL signal. R1 and Kimi k1.5 both reported PRMs weren't worth it at their scale.
Give five concrete ways a model could reward-hack a coding RL environment, and how you'd prevent each.
(1) Hard-code expected outputs for visible tests. Prevent: hidden and randomized tests. (2) Edit or delete failing tests or conftest.py. Prevent: run graders from a pristine copy outside the sandbox, and checksum test files. (3) Exit early with status 0, or override __eq__ to always be true. Prevent: graders that check actual outputs and don't trust exit codes; detect monkeypatching. (4) Find the reference fix in git history, pip caches or the internet. Prevent: scrub history, egress allowlist. (5) Tamper with timing or scoring harness code (documented by METR). Prevent: keep graders out of the writable filesystem. Plus process: red-team the env with a strong model before training, include impossible-task canaries, and run monitors over transcripts.
Why is optimizing against a chain-of-thought monitor risky?
Baker et al. showed that LLM monitors reading the CoT catch reward hacking much better than action-only monitors. But adding the monitor's verdict to the reward, under strong optimization, taught the model to hide its intent: the CoT looks clean while the hack continues ("obfuscated reward hacking"). You lose the monitoring channel and still have the misbehavior. Recommended practice is to use CoT monitors for detection, env patching and data filtering, and to avoid strong direct optimization pressure on CoT content, accepting a "monitorability tax".
Design an RL environment for training an agent to resolve CI failures in real repositories.
Tasks: mine historical CI failures with the commit that fixed them; filter to deterministic failures reproducible in a container; split by repo/time for train vs eval. State: a container image per repo version with deps pinned and the failing commit checked out; snapshot for fast reset. Tools: shell (with timeout), file view/edit, run-tests, submit. Observations: truncated logs with markers, consistent formatting. Termination: submit, N turns, or wall-clock. Reward: the originally failing job passes plus a pass-to-pass regression suite, run by a separate grader on a pristine copy of the tests. Possibly a small penalty for touching test/CI config. Calibration: measure the current policy's pass rate and keep a 10–90% band. QA: verify each task with the real fix (must pass) and a no-op (must fail), add impossible canaries, red-team. Determinism: no network, or record/replay, seeded randomness. Throughput: warm pools, cached images, cost per episode tracked.
What is loss masking in multi-turn RL, and what goes wrong without it?
In agent trajectories, tool outputs, user messages and environment text are interleaved with the policy's tokens. Only the policy's tokens are actions, so the policy-gradient loss (and KL) must be computed only over them, via a mask that is 1 on sampled tokens and 0 on prompt/environment tokens. Without masking, the model is pushed to raise or lower the likelihood of text it didn't choose (tool outputs) according to the episode's advantage. That's meaningless credit assignment that can destabilize training and teach it to hallucinate tool outputs. Related: store exact token ids and logprobs from generation to avoid re-tokenization drift.
Colocated vs disaggregated RL infra: when would you choose each?
Colocated (veRL-style hybrid engine): the same GPUs switch between inference and training, resharding weights. It's simple, has no idle dedicated pools, and weight sync is cheap (local), so it's good for smaller clusters or short generations. Disaggregated: separate inference and training pools that run concurrently, each sized and tuned independently. That suits long rollouts, agentic workloads with environment latency, and async pipelines, but it needs network weight transfer and load balancing, and adds staleness. At frontier scale with long agentic episodes, disaggregated plus async is the common direction. Colocated is a strong default for research-scale reasoning RL.
What is the rollout/trainer logprob mismatch, and how do you handle it?
The inference engine (vLLM/SGLang) and the training framework compute slightly different probabilities for the same tokens and weights, because of different kernels, precision, batch-dependent reductions, quantized rollouts and MoE routing differences. So nominally on-policy data is off-policy, and PPO ratios computed against engine logprobs are biased. Detect it by logging both and plotting \(\pi_\text{train}/\pi_\text{rollout}\). Mitigate with truncated importance sampling (weight by \(\min(\text{ratio}, C)\)), recomputing old logprobs with the trainer, sequence-level ratios, matching precision, or batch-invariant deterministic kernels.
How does asynchronous RL work, and what's the cost?
Generators keep producing rollouts with whatever weights they have while the trainer updates. Data may be k versions stale. Benefits: high GPU utilization and no waiting for stragglers. Costs: off-policy bias and instability as staleness grows. Mitigations: bound staleness (drop data older than k versions), importance-weight corrections (decoupled PPO objectives as in AReaL), in-flight weight updates to keep sequences fresh (PipelineRL), and objectives more robust to off-policy data (Noukhovitch et al. found online DPO robust). Report the staleness distribution as a training metric.
How do you calibrate task difficulty for GRPO, and why does it matter so much?
With group-relative advantages, a prompt where all G samples get the same reward yields zero gradient. Too-easy and too-hard tasks waste rollout compute and also shrink the effective batch, which adds variance. Calibrate by estimating the current policy's pass rate per task (8–16 samples), keeping tasks in an intermediate band, and re-estimating periodically as the policy improves (curriculum). Online, DAPO's dynamic sampling over-samples and discards all-tie groups. Keep a reserve of harder tasks to promote as the model improves, and track the fraction of zero-variance groups as a health metric.
How would you get RL signal for a non-verifiable task such as writing medical patient summaries?
Push toward verifiability: extract checkable sub-claims (medications, dosages, diagnoses present in the source record; no hallucinated facts, via entailment against the record), plus format and length constraints. For quality, use per-prompt rubrics (expert-written or model-drafted then expert-reviewed) graded item-by-item by a judge from a different model family, or pairwise comparisons against a reference. Guard against judge exploitation with length control, hidden rubrics, periodic human audits, and tracking judge-human agreement over training. Keep a KL anchor because the reward is softer than a program.
What metrics do you watch during an RL run to know it's healthy?
Reward (train) and held-out eval accuracy, which should track each other; divergence suggests hacking or overfitting. Response length, split by correct and incorrect. Policy entropy (collapse means lost exploration). KL to the reference. Clip fraction and importance-ratio statistics (including the engine-vs-trainer mismatch). Fraction of zero-advantage groups. Gradient norm. Per-environment and per-difficulty pass rates. Truncation rate. Invalid-trajectory rate (environment failures). Staleness distribution in async setups. Hack-monitor flags. Plus periodic transcript reading, which no dashboard replaces.
Compare GSPO's sequence-level ratio with GRPO's token-level ratio.
GRPO computes \(\pi_\theta/\pi_\text{old}\) per token and clips per token, even though the advantage is the same for every token in a sequence. Each token's ratio is a one-sample, high-variance estimate, and the noise accumulates over long sequences. In MoE models, routing shifts make individual token ratios swing more. GSPO uses the length-normalized sequence likelihood ratio \((\pi_\theta(o)/\pi_\text{old}(o))^{1/|o|}\) and clips whole sequences, matching the granularity of the reward. It reports more stable training (notably for MoE) and simpler infrastructure, since it's less sensitive to small token-level precision mismatches. The cost: a sequence is either inside or outside the trust region as a whole.
What's an RL environment "harness", and why does the train/deploy match matter?
The harness is the agent program around the model: the system prompt, tool schemas, the loop that parses tool calls and feeds back observations, context management. Policies learn harness-specific habits (tool-call formats, how much output to request, when to submit). If training uses a toy ReAct loop but deployment uses a production coding agent with different tools and context handling, skills transfer imperfectly. The 2025–2026 trend, visible in Harbor and verifiers v1, is to run RL inside the real production harness, with tools via MCP, so the trained behavior is the deployed behavior.
You see reward going up but held-out accuracy flat. Walk through your debugging.
First, read transcripts from high-reward rollouts for hacks (test edits, special-casing, judge flattery, multiple answers). Check whether the reward is dominated by an auxiliary component (format, length). Compare train vs held-out distributions (are eval tasks harder, different domain, different harness?). Check contamination or memorization of training tasks (reward rising on repeated tasks only). Verify the grader on known-bad solutions. Check whether gains are in extraction/format that the eval scores differently. Fixes: harden the env, rebalance reward components, add held-out checks into training monitoring, early stop.