Post-training & Alignment
Pretraining gives you a model that has read the internet and can continue any document. Post-training turns that model into an assistant: one that follows instructions, holds a conversation, reasons step by step, uses tools, declines harmful requests and doesn't refuse harmless ones. This page covers the whole stack: supervised fine-tuning, reward models, RLHF with PPO, DPO and its relatives, RL with verifiable rewards, AI feedback, parameter-efficient fine-tuning (LoRA/QLoRA), distillation, model merging and safety training. Interviewers ask about it because nearly every applied AI role touches some part of it. You might be fine-tuning an open model, building a preference dataset, choosing between LoRA and full fine-tuning, or debugging a model that got worse after tuning. And the theory (the KL-regularized objective, the DPO derivation) is a favourite whiteboard topic.
TL;DR: the 8–12 things to be able to say out loud
- Base vs instruct: pretraining puts in the knowledge. Post-training mostly shapes format, style, behaviour and which capabilities get used. A small amount of high-quality SFT data goes a long way (LIMA), but RL-style training adds real capability on reasoning tasks, so "alignment is superficial" is only partly true.
- Modern pipeline: SFT → preference optimization (DPO or RM+PPO) → RL with verifiable rewards (RLVR, usually GRPO-family) → safety and general-alignment tuning, often in iterative rounds where each round's model generates the next round's data.
- SFT is next-token cross-entropy on (prompt, response) pairs, with the loss masked on prompt tokens, using a fixed chat template. Packing gives throughput. Data quality and diversity beat volume.
- Reward model: a scalar head trained with the Bradley–Terry loss \(-\log\sigma(r(x,y_w)-r(x,y_l))\) on pairwise preferences. It is a proxy, and optimizing it too hard brings Goodhart effects (RM overoptimization).
- RLHF objective: maximize \(\mathbb{E}[r(x,y)] - \beta\,\mathrm{KL}(\pi_\theta\,\|\,\pi_\text{ref})\). PPO needs four models (policy, reference, reward, value), online generation and many hyperparameters, so it is expensive and finicky.
- DPO: the KL-regularized objective has a closed-form optimum \(\pi^*\propto\pi_\text{ref}\,e^{r/\beta}\). Invert it to write the reward in terms of the policy, plug that into Bradley–Terry, and the partition function cancels. You get a classification loss on preference pairs: no RM, no sampling, two models.
- Variants: IPO (bounded loss, less overfitting), KTO (unpaired thumbs-up/down data), ORPO (one stage, no reference model), SimPO (length-normalized, no reference). Online/iterative DPO closes much of the gap to PPO.
- Online vs offline: on-policy data generally wins for the final model. Offline DPO wins on simplicity and cost. Most frontier recipes use both.
- RLVR / GRPO: for math and code, replace the learned RM with a programmatic checker. GRPO drops the critic and uses group-normalized rewards as advantages. This is what produced the reasoning models (DeepSeek-R1); see A5.
- LoRA: freeze \(W_0\) and learn \(\Delta W = \tfrac{\alpha}{r}BA\) with rank \(r\ll d\). That cuts trainable parameters and optimizer memory by orders of magnitude, and the update merges back so inference costs nothing extra. QLoRA = 4-bit NF4 frozen base + LoRA. LoRA "learns less and forgets less" than full fine-tuning.
- Distillation (logit-level KD, or SFT on teacher outputs) is how most small models get strong, especially reasoning models. Merging (task arithmetic, TIES, SLERP) combines fine-tunes without extra training.
- Safety tuning trades harmlessness against over-refusal. Safety alignment is shallow and easily undone by fine-tuning, and jailbreak robustness is now handled by defence in depth (training plus classifiers), not by the model alone.
Why post-train? Base models vs assistants
A pretrained ("base") model is a document completer. Ask it "What is the capital of France?" and it may answer, or it may continue with four more quiz questions, because that is what such text looks like on the web. The base model already holds the knowledge and much of the skill. What it lacks is a consistent persona and interface: knowing that a user turn ends and an assistant turn begins, that it should answer helpfully and stop, that it should follow formatting instructions, and that some requests should be declined.
The canonical demonstration was InstructGPT: humans preferred outputs from a 1.3B-parameter model fine-tuned with SFT and RLHF over outputs from the 175B GPT-3 base model, despite the 100× size gap. Ouyang+ 2022 Post-training is extremely leveraged: it uses a tiny fraction of pretraining compute (historically well under ~1–2%, though that share has grown a lot with RL-heavy reasoning training) and changes perceived quality dramatically.
The superficial alignment hypothesis
LIMA fine-tuned a 65B LLaMA on just 1,000 carefully curated prompt–response pairs, with no RLHF. Its responses were judged equal or preferred to GPT-4's in a substantial fraction of cases. Zhou+ 2023 The authors proposed the superficial alignment hypothesis: almost all knowledge and capability is learned in pretraining, and alignment mostly teaches which sub-distribution of formats and styles to use when talking to users. Supporting evidence: URIAL showed that base models aligned purely in context (a few stylistic examples plus a system prompt, no weight updates) get surprisingly close to their SFT+RLHF counterparts. Token-distribution analysis showed that aligned and base models disagree mainly on stylistic tokens ("Sure", "However", discourse markers), not on content tokens. Lin+ 2023
Where the hypothesis breaks
- Reasoning via RL. Large-scale RL on verifiable tasks (DeepSeek-R1, OpenAI o-series) produces large gains on math and code benchmarks, long chains of thought, self-verification and backtracking. DeepSeek-AI 2025 That is more than format. Whether it is new capability or better elicitation of latent capability is debated: one study found RLVR raises pass@1 but the base model catches up and even overtakes at large pass@k, suggesting RL mostly sharpens the distribution towards solutions the base model could already sample. Yue+ 2025 Treat this as contested.
- Long-tail skills and robustness. Tool use, long-horizon agentic behaviour, calibrated refusals, multi-turn consistency and instruction hierarchy all need far more than 1,000 examples. Production post-training sets now run to millions of examples plus RL environments.
- Small data also elicits reasoning. In the same spirit as LIMA, LIMO showed a few hundred curated long-reasoning examples can unlock strong math reasoning in a capable base model. Ye+ 2025 The lesson isn't "data doesn't matter". It is "the base model sets the ceiling, and post-training decides how much of it you reach".
Think of pretraining as building a huge library of latent "characters" and skills, and post-training as selecting and sharpening one character (the helpful assistant), then drilling specific skills on top. SFT selects the character. Preference tuning polishes its taste. RL with verifiable rewards drills skills where you can check the answer.
"Is alignment superficial?" is a judgement question. A strong answer cites LIMA/URIAL as evidence that style and format are cheap to change. It then names the counter-evidence: RL-driven reasoning gains, tool use and robustness all need much more data and compute. It closes with the open question of whether RL creates capability or elicits it (pass@k evidence). Avoid one-sided answers.
- LIMA (Zhou+ 2023): the 1,000-example paper and the superficial alignment hypothesis.
- URIAL / Unlocking Spell (Lin+ 2023): token-shift analysis and in-context alignment.
- InstructGPT (Ouyang+ 2022): the original three-step SFT → RM → PPO recipe.
The modern post-training pipeline, end to end
The 2022 InstructGPT recipe was three steps: (1) SFT on human demonstrations, (2) train a reward model on human rankings, (3) optimize the policy against the RM with PPO. Today's recipes keep the skeleton but differ in four main ways. Preference learning is often done with DPO-family losses instead of PPO. A verifiable-reward RL stage targets math, code and instruction-following constraints. Much of the data is synthetic (generated by the model itself or a stronger teacher, then filtered). And the whole thing runs in iterative rounds.
Published recipes worth knowing
| Model | Recipe (as published) | Notable choices |
|---|---|---|
| Llama 2-Chat (2023) | SFT → separate helpfulness and safety RMs → iterative rejection-sampling fine-tuning + PPO across several versions Touvron+ 2023 | Two RMs to manage the helpfulness/safety tension. Margin term in the BT loss based on how strongly annotators preferred one response. "Ghost Attention" for system-prompt adherence across turns. |
| Llama 3 (2024) | Several rounds (six) of: train RM → rejection sampling with the RM → SFT → DPO. Llama Team 2024 | Chose DPO over PPO for stability and scale. Heavy synthetic data for code, math, tool use and long context. Model averaging across runs at each stage. |
| Tülu 3 (AI2, 2024) | SFT → DPO (on-policy-ish preference data: completions from the SFT model and others, rated by an LLM judge) → RLVR (PPO with rewards only when a verifier says the answer is correct). Lambert+ 2024 | Fully open data, code and evals. Popularized the term "RLVR". Decontamination of training prompts against evals. |
| DeepSeek-R1 (2025) | R1-Zero: pure GRPO on the base model with rule-based accuracy + format rewards, no SFT. R1: small long-CoT "cold start" SFT → reasoning RL (plus a language-consistency reward) → rejection-sample ~800k examples (reasoning + general) for SFT → final RL over all scenarios. Then distil into Qwen/Llama models by SFT alone. DeepSeek-AI 2025 | Showed long CoT and "aha" self-correction emerging from outcome-only RL. Reported that process reward models and MCTS did not pay off at scale for them. Distilled small models beat RL-trained small models. |
| Qwen2.5 (2024) | Large SFT (>1M examples) → offline RL (DPO) → online RL (GRPO) against an RM. Qwen Team 2024 | Explicit "offline then online" two-stage RL. |
| Qwen3 (2025) | Four stages for flagships: long-CoT cold start → reasoning RL → "thinking mode fusion" (SFT mixing thinking and non-thinking data so one model can do both) → general RL. Smaller models get strong-to-weak distillation from the flagships. Qwen Team 2025 | One model with switchable thinking and non-thinking modes. Distillation is much cheaper than running the full pipeline on every size. |
| OLMo 3 (AI2, late 2025) | SFT → DPO → scaled-up RLVR, for both "Think" and "Instruct" variants, with open data. OLMo Team 2025 | DPO pairs chosen for the contrast between responses ("delta learning"), even when the chosen response alone wouldn't be a good SFT target. |
Frontier labs (OpenAI, Anthropic, Google) don't publish full recipes, and open recipes change every few months. As of late 2026 the broad consensus is: SFT → cheap offline preference stage → large-scale online RL (verifiable rewards plus learned/rubric rewards) → distillation into smaller models. The share of compute going to RL has been growing fast. Check the latest technical reports (Qwen, DeepSeek, Llama, OLMo, Kimi, Nemotron) before quoting specifics.
Why iterative rounds?
Preference and RL data go stale. A reward model trained on outputs of model v1 is less accurate on outputs of v2, whose distribution has shifted (and v2 may have found the RM's blind spots). Each round, labs therefore (a) sample fresh responses from the current best model, (b) collect new preferences or judge scores on those, (c) retrain or refresh the RM, and (d) run SFT/DPO/RL again. Rejection sampling (best-of-N according to an RM or verifier, then SFT on the winners) is the cheapest way to turn a better RM into a better policy, which is why Llama 2/3 lean on it. This is also why the "offline vs online" distinction blurs: iterative DPO with fresh on-policy samples each round is "semi-online".
"Walk me through how you'd post-train an open base model into a chat assistant." A strong answer covers: data (prompt sources, synthetic generation, decontamination), SFT with a fixed chat template and prompt masking, a preference stage (DPO first because it's cheap), RLVR for math, code and format constraints, a safety pass with over-refusal evals, then iterate. Mention evals at every stage (A8) and watching for regressions in the capabilities you didn't train on.
- Tülu 3 report: the most detailed fully open recipe, including data mixes and ablations.
- Llama 3 Herd, §4 (post-training): rejection sampling + DPO at scale.
- RLHF book (Lambert): free, continuously updated textbook on the whole pipeline.
Supervised fine-tuning (SFT)
SFT is ordinary next-token prediction on curated (prompt, response) conversations. Three engineering details trip people up: the chat template, loss masking and packing. Above all of them sits the question of data.
Chat templates
A chat model sees a conversation serialized into a single token stream with special tokens marking roles and turn boundaries. The exact template is model-specific, and using the wrong one at inference silently degrades quality. Libraries ship templates with the tokenizer (for example Jinja templates in Hugging Face tokenizer.apply_chat_template). HF docs A schematic ChatML-style example:
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
What's 17 * 23?<|im_end|>
<|im_start|>assistant
17 * 23 = 391.<|im_end|>
The end-of-turn token matters: the model must learn to emit it, or generation never stops. A common bug is a template that never puts the EOS/end-of-turn token in the loss.
Loss masking on prompt tokens
The SFT loss is computed only on assistant tokens:
$$\mathcal{L}_\text{SFT}(\theta) = -\sum_{t \in \text{assistant}} \log \pi_\theta(y_t \mid x, y_{<t})$$System and user tokens are still in the context (the model attends to them) but get label -100 (ignored). Why: (1) you want the model to learn to respond, not to imitate users, (2) prompt tokens are often noisy, repetitive or adversarial, (3) without masking, long prompts dominate the gradient. In multi-turn data you typically train on every assistant turn, or only the final one if earlier turns came from a weaker model. Some work finds training on prompts slightly helps when responses are very short, so it is a tunable choice, but masking is the default.
Packing
SFT examples vary a lot in length. Padding each batch to the longest example wastes compute (often 30–50% of tokens on short-heavy chat data). Packing concatenates several examples into one fixed-length sequence. Done naively, tokens in example B can attend to example A ("cross-contamination"). The correct approach resets position IDs per example and uses a block-diagonal attention mask, implemented efficiently with variable-length attention kernels (for example FlashAttention's varlen interface). Krell+ 2021 Note that packing changes the effective per-example loss weighting: summing over tokens across a packed batch weights long examples more than a per-example mean would. Some frameworks therefore normalize per example.
Data quality over quantity
- Quality and diversity dominate. LIMA's 1,000 examples were hand-picked for diversity and a consistent "helpful assistant" style. Thousands of near-duplicates add little. Deduplicate (exact plus near-duplicate via MinHash or embeddings) and cluster prompts to balance topics.
- Consistency of style. Mixed styles (some terse, some verbose, some with markdown) produce an inconsistent model. Many teams rewrite all responses into one house style.
- Correctness. SFT faithfully imitates errors. One wrong-but-confident answer pattern, repeated, teaches hallucination. Verify math and code answers by execution where possible.
- Don't teach knowledge the model lacks. Fine-tuning on facts the base model doesn't know can increase hallucination: the model learns to produce confident answers it can't ground. This is one reason SFT data is often generated by the model itself plus filtering rather than written by external experts.
- Decontamination. Remove any prompt that overlaps with evaluation sets (n-gram overlap). Tülu 3 found and removed contamination in several popular public datasets. Lambert+ 2024
Synthetic data generation
Most SFT data is now synthetic. The main patterns:
| Method | How it works | Use |
|---|---|---|
| Self-Instruct | Start from 175 seed tasks. Prompt the model to generate new instructions, inputs and outputs. Filter by ROUGE overlap and heuristics. Iterate. Wang+ 2022 | Bootstrapping instruction diversity (the basis of Stanford Alpaca). |
| Evol-Instruct | An LLM rewrites seed instructions to be harder ("in-depth": add constraints, deepen, add reasoning steps) or different ("in-breadth": new topics), with elimination of failed evolutions. Xu+ 2023 | Raising difficulty and complexity (WizardLM, WizardCoder). |
| Magpie | Feed an aligned model only the chat-template prefix up to the user turn. It "autocompletes" a plausible user query, then answers it. Xu+ 2024 | Large-scale prompt harvesting without seeds. |
| Distillation from a teacher | Answer real or synthetic prompts with a stronger model (Orca used explanation traces from GPT-4). Mukherjee+ 2023 | Most open chat models, reasoning distillation. Check the teacher's license terms. |
| Rejection sampling (RFT / best-of-N) | Sample K responses per prompt from the current model. Keep those an RM ranks highest, or those a verifier marks correct. SFT on them. Yuan+ 2023 | Self-improvement loop for math and code, and Llama-style iterative rounds. |
| STaR | Generate rationales. Keep those reaching the correct answer. For failures, "rationalize" given the answer hint. Fine-tune. Repeat. Zelikman+ 2022 | Early self-taught reasoning, an ancestor of RLVR. |
Saying "SFT on synthetic data causes model collapse". Recursive training on unfiltered generations can lose diversity and the tails of the distribution. But filtered, verified, teacher-generated or rejection-sampled synthetic data is the backbone of every modern recipe. The key ingredients are a filter (verifier, RM, judge) and mixing with fresh, diverse prompts.
SFT hyperparameters and magnitudes
Typical full fine-tuning SFT uses a learning rate around 1e-5 to 2e-5 for 7–70B models (much lower than pretraining peak LR), 1–3 epochs (more epochs overfit and hurt diversity), cosine or linear decay with short warmup, and global batch sizes of tens to a few hundred sequences. LoRA uses learning rates roughly 10× higher (see below). Watch for loss dropping very low on the train set, which usually means memorization of a small dataset rather than better behaviour.
Expect "Why mask the prompt?", "How does packing work without contaminating attention?", "Your fine-tuned model never stops generating. Why?" (EOS not trained or template mismatch), and "How would you generate SFT data for domain X?" A strong answer for the last one: seed prompts from real usage, synthetic expansion (Evol/Magpie), teacher or self-generated responses, automated filtering (verifiers, judges, dedup), human spot-checks, decontamination, and a held-out eval built before training.
- TRL documentation: reference implementations of SFTTrainer (packing, completion-only loss), DPO, KTO, ORPO, GRPO.
- WizardLM / Evol-Instruct: the template for difficulty-raising synthetic data.
- Magpie: seedless prompt synthesis from aligned models.
Reward models
A reward model (RM) turns human (or AI) judgements into a differentiable scalar you can optimize. The standard architecture is the SFT model with its LM head replaced by a linear head that outputs one number \(r_\phi(x,y)\), read at the last token.
Pairwise data and the Bradley–Terry loss
Humans are bad at absolute scores ("rate this 1–10") and much more consistent at comparisons ("which is better?"). So data is collected as pairs: for a prompt \(x\), show two (or more) responses and record which is preferred, \(y_w \succ y_l\). The Bradley–Terry model assumes each response has a latent quality score and
$$P(y_w \succ y_l \mid x) = \sigma\big(r(x,y_w) - r(x,y_l)\big) = \frac{e^{r(x,y_w)}}{e^{r(x,y_w)} + e^{r(x,y_l)}}$$Training maximizes the likelihood of the observed preferences:
$$\mathcal{L}_\text{RM}(\phi) = -\,\mathbb{E}_{(x,y_w,y_l)}\Big[\log \sigma\big(r_\phi(x,y_w) - r_\phi(x,y_l)\big)\Big]$$Only differences matter, so the RM is defined up to a per-prompt constant (people often normalize rewards before RL). With \(K\) ranked responses you get \(\binom{K}{2}\) pairs per prompt. InstructGPT ranked 4–9 responses and batched all pairs from one prompt together to avoid overfitting. Ouyang+ 2022 Llama 2 added a margin \(m(r)\) inside the sigmoid for "significantly better" pairs. Touvron+ 2023 RMs typically reach only ~65–75% agreement with held-out human preferences on hard chat data. Inter-annotator agreement itself is often in the same range, so preference data is noisy by nature.
Collecting preference data in practice
- Sampling responses: use the current policy (on-policy) and a mix of other models for diversity. Vary temperature. Responses that are too similar give low-information pairs.
- Guidelines: annotators need a written rubric (helpfulness, honesty, harmlessness and their priority order). Most disagreement comes from unclear guidelines, not careless raters.
- Multi-attribute: separate RMs or heads for helpfulness and safety (Llama 2) avoid a single scalar that has to trade them off implicitly.
- Cost: expert comparisons cost dollars each. AI feedback (see RLAIF) costs fractions of a cent, which is why most open preference sets (for example UltraFeedback-style) are LLM-judged.
Reward overoptimization (Goodhart's law)
The RM is a proxy fitted on a finite, off-policy dataset. A policy optimized hard against it finds inputs where the proxy is wrong: verbosity, sycophancy, confident tone, lists and headers, flattering the user, or outright gibberish that the RM happens to score highly. Gao et al. measured this with a synthetic setup: a large "gold" RM labels data, a smaller proxy RM is trained on those labels, and a policy is optimized against the proxy. As KL distance from the initial policy grows, proxy reward keeps rising while gold reward rises, peaks, then falls. Gao+ 2022 They fit functional forms in \(d=\sqrt{\mathrm{KL}(\pi\|\pi_\text{init})}\): roughly \(R_\text{bon}(d)=d(\alpha-\beta d)\) for best-of-N and \(R_\text{RL}(d)=d(\alpha-\beta\log d)\) for RL. Larger RMs and more RM data push the peak further out.
A well-documented specific case is length bias. RMs tend to prefer longer answers, and a large share of measured RLHF "improvement" can come from length alone. Singhal+ 2023 This is why AlpacaEval introduced a length-controlled win rate. Dubois+ 2024 Mitigations: KL penalty, RM ensembles, length penalties or normalization, refreshing the RM on on-policy samples, early stopping on a held-out gold metric, and constraining the output format.
Outcome vs process reward models
Outcome RM (ORM)
Scores the final answer, or the whole response. Cheap labels: for math, just check the answer. Sparse credit assignment: a long solution with one bad step and a lucky right answer gets full credit. A rule-based verifier is the limiting case of an ORM.
Process RM (PRM)
Scores each reasoning step. "Let's Verify Step by Step" collected ~800k human step-level labels (PRM800K). Their PRM beat an ORM for best-of-N selection on MATH. Lightman+ 2023 Labels are expensive. Math-Shepherd derived step labels automatically via Monte Carlo rollouts from each step. Wang+ 2023
In practice PRMs have proven most useful for search and reranking at inference time. For RL training at scale, DeepSeek-R1 reported PRMs were hard to define, label and protect from reward hacking, and used outcome rewards instead. DeepSeek-AI 2025 Work on PRM pitfalls (for example noisy Monte Carlo labels) continues. Zhang+ 2025 Implicit PRMs, which derive step rewards from an outcome-trained model, are an active direction. Cui+ 2025
Generative reward models and LLM-as-judge
Instead of a scalar head, ask an LLM to judge in natural language: "Which response is better and why?", or "Is this solution correct? Answer yes/no", and use the probability of "yes" as the reward. Generative verifiers trained with next-token prediction (optionally reasoning with CoT before the verdict) outperform discriminative RMs on reasoning tasks and can spend more inference compute to judge better. Zhang+ 2024 Liu+ 2025 LLM judges correlate well with human preference on chat (MT-Bench reported >80% agreement with humans for GPT-4, similar to human–human agreement) but have known biases: position, verbosity and self-preference. Zheng+ 2023 A 2025 trend extends RL beyond verifiable domains by having a judge score responses against instance-specific rubrics ("rubrics as rewards"). Gunjal+ 2025 OpenAI described "rule-based rewards" for safety behaviour: an LLM grades responses against explicit propositions ("refuses without judgemental language") instead of holistic preference. Mu+ 2024
Classic probes: "Write the RM loss", "Why pairwise instead of absolute scores?", "What is reward hacking? Give an example", "PRM vs ORM?", "How do you know your RM is good?" For the last one: held-out preference accuracy, RewardBench-style benchmarks Lambert+ 2024, correlation with downstream policy quality (more predictive but expensive), and checks for length and style bias. A senior answer notes that RM benchmark accuracy does not always predict how good a policy you get from RL against it.
- Scaling Laws for RM Overoptimization (Gao+ 2022): the gold-vs-proxy curves.
- Lilian Weng: Reward Hacking in RL: broad survey of reward hacking, including in LLMs.
- Let's Verify Step by Step: PRMs vs ORMs with careful human labels.
RLHF with PPO
The KL-regularized objective
Everything in preference optimization starts from this objective:
$$\max_{\pi_\theta}\; \mathbb{E}_{x\sim\mathcal{D},\,y\sim\pi_\theta(\cdot|x)}\big[r_\phi(x,y)\big] \;-\; \beta\, \mathbb{D}_\text{KL}\big(\pi_\theta(\cdot|x)\,\|\,\pi_\text{ref}(\cdot|x)\big)$$\(\pi_\text{ref}\) is the frozen SFT model. The KL term does three jobs. It keeps the policy where the RM is accurate, which limits overoptimization. It preserves the fluency and diversity of the SFT model, which prevents mode collapse onto one high-reward template. And it is a regularizer against forgetting. \(\beta\) sets the trade-off: small \(\beta\) means more reward and more hacking, large \(\beta\) means the policy barely moves. Typical values are around 0.01–0.1, sometimes adapted to hit a target KL.
In implementation the KL is applied per token as a reward-shaping term. The RM score arrives at the last token:
$$r_t = -\beta\,\log\frac{\pi_\theta(y_t|x,y_{<t})}{\pi_\text{ref}(y_t|x,y_{<t})} \;+\; \mathbb{1}[t=T]\; r_\phi(x,y)$$Generation is treated as an episodic MDP: state = prompt + tokens so far, action = next token, episode ends at EOS.
PPO in one paragraph
PPO is a policy-gradient method. Its gradient pushes up the log-probability of actions with positive advantage \(A_t\) (better than expected) and down for negative ones. To reuse a batch of samples for several gradient steps without the policy drifting too far, it clips the importance ratio \(\rho_t=\pi_\theta/\pi_{\theta_\text{old}}\): Schulman+ 2017
$$\mathcal{L}^\text{CLIP} = \mathbb{E}_t\Big[\min\big(\rho_t A_t,\; \mathrm{clip}(\rho_t, 1-\epsilon, 1+\epsilon)\,A_t\big)\Big]$$Advantages come from a learned value function (critic) \(V_\psi(s_t)\) via Generalized Advantage Estimation (GAE), which blends multi-step returns to balance bias and variance. The critic is usually initialized from the RM or SFT model and trained with a regression loss on returns.
Why it's expensive and unstable
- Memory: four model copies. For a 7B model with mixed-precision Adam (~16 bytes/param for weights + grads + optimizer state), actor ≈ 112 GB and critic ≈ 112 GB, plus two frozen bf16 copies (≈14 GB each). That is roughly 250 GB before activations and the KV cache for generation, so even a 7B run needs multi-GPU sharding. Tricks: LoRA adapters on a shared base (the reference is "base with adapters off"), sharing actor/critic trunks, offloading the frozen models.
- Generation is the bottleneck: autoregressive sampling of long responses dominates wall-clock time. Modern systems run a separate inference engine (vLLM/SGLang-style) for rollouts and sync weights to the trainer. That brings off-policy lag and extra engineering (see A5).
- Many interacting hyperparameters: β/KL target, clip ε, GAE λ and γ, value-loss coefficient, LR for actor and critic, number of PPO epochs, batch and minibatch sizes, reward normalization and whitening, sampling temperature. Small implementation details (advantage normalization, reward clipping, EOS handling) matter a lot.
- Critic difficulty: predicting expected final reward from partial text is hard. A bad critic gives noisy advantages, and noisy advantages give unstable updates.
- Reward hacking: the online loop actively searches for RM exploits. KL can collapse (policy barely moves) or explode (policy diverges into degenerate text).
PPO-RLHF is "generate, grade, nudge". The RM grades whole answers. The critic guesses, mid-answer, how the grade will turn out, so you know which tokens deserve credit. The reference model is a leash. Each moving part is a source of variance, which is why the field keeps inventing methods that remove parts: DPO removes the RM, critic and sampling. GRPO removes the critic. RLVR replaces the RM with a checker.
"Name the four models and say which are trained." "Why do we need the KL term?" (stay in the RM's trust region, avoid mode collapse and hacking, preserve capabilities). "What happens if β=0?" (reward hacking and degenerate text). "Why is PPO-RLHF expensive?" (memory, generation-bound, sensitive to hyperparameters). Bonus: KL is a reverse KL \(\mathrm{KL}(\pi_\theta\|\pi_\text{ref})\), which is mode-seeking: the policy may drop modes of the reference but is penalized for putting mass where the reference has little.
- HF blog: Illustrating RLHF: friendly visual walkthrough.
- Chip Huyen: RLHF: engineering-oriented explanation of the pipeline.
- Learning to summarize from human feedback (Stiennon+ 2020): the first convincing large-scale RLHF-on-LM paper.
Direct Preference Optimization (DPO)
DPO removes the reward model and the RL loop entirely. Rafailov+ 2023 The derivation is short and a favourite interview question.
Derivation intuition (three steps)
Step 1 is a standard result: maximizing \(\mathbb{E}_\pi[r] - \beta\mathrm{KL}(\pi\|\pi_\text{ref})\) over distributions gives a Gibbs/Boltzmann reweighting of the reference. Z(x) is intractable (a sum over all possible sequences), which is why you can't just compute \(\pi^*\). The trick in step 3 is that you never need it. The resulting loss:
$$\mathcal{L}_\text{DPO}(\theta) = -\,\mathbb{E}_{(x,y_w,y_l)}\left[\log\sigma\left(\beta\log\frac{\pi_\theta(y_w|x)}{\pi_\text{ref}(y_w|x)} - \beta\log\frac{\pi_\theta(y_l|x)}{\pi_\text{ref}(y_l|x)}\right)\right]$$The quantity \(\hat r_\theta(x,y)=\beta\log\frac{\pi_\theta(y|x)}{\pi_\text{ref}(y|x)}\) is the implicit reward. That is the paper's subtitle: "your language model is secretly a reward model". The log-probability of a sequence is just the sum of token log-probs, so the loss needs four forward passes per pair (policy and reference on chosen and rejected). The reference log-probs can be precomputed once.
What the gradient does
$$\nabla_\theta\mathcal{L}_\text{DPO} = -\beta\,\mathbb{E}\Big[\underbrace{\sigma\big(\hat r_\theta(x,y_l)-\hat r_\theta(x,y_w)\big)}_{\text{weight: how wrong we are}}\big(\nabla\log\pi_\theta(y_w|x) - \nabla\log\pi_\theta(y_l|x)\big)\Big]$$It raises the likelihood of chosen and lowers rejected, weighted by how badly the implicit reward currently mis-orders the pair. Pairs already ranked correctly by a wide margin contribute almost nothing. That weighting is what stops it degenerating like naive "unlikelihood" training.
The role of β
β is the same KL-strength parameter as in RLHF. Small β (for example 0.01–0.05) lets the policy move far from the reference and fit preferences aggressively, with more risk of overfitting and verbosity. Large β (0.3–0.5) keeps it close. β = 0.1 is the common default. In DPO, β also scales the margin at which the sigmoid saturates, so it interacts with learning rate. DPO learning rates are small (around 5e-7 to 5e-6 for full fine-tuning), and 1–3 epochs is typical. It overfits quickly.
Known failure modes
- Both likelihoods go down. DPO only constrains the margin. Commonly the log-prob of the chosen response also decreases, just less than the rejected one, and probability mass moves to responses that are in neither. Remedies: add an SFT/NLL term on chosen responses (as in RPO-style and Llama 3 variants), or use a higher β.
- Length exploitation. Chosen responses are often longer, and DPO learns "longer = better". Length-regularized DPO explicitly penalizes this. Park+ 2024
- Off-policy mismatch. If the preference pairs come from other models, the policy is learning to rank text it would never generate. The implicit reward generalizes poorly to its own outputs. This is the core argument for online/iterative variants.
- Overfitting with deterministic preferences. If \(y_w\) always beats \(y_l\), the BT likelihood is maximized by pushing the margin to infinity. The KL regularization effectively disappears because there is no explicit KL term in the loss. This motivates IPO.
Variants
| Method | Key change | Loss sketch | Why use it |
|---|---|---|---|
| IPO Azar+ 2023 | Replaces log-sigmoid with a squared loss towards a fixed target margin. Part of a general "ΨPO" framework. | \(\big(h_\theta(x,y_w,y_l) - \tfrac{1}{2\tau}\big)^2\), where \(h\) is the log-ratio difference | Bounded target, so it doesn't push margins to infinity and stays regularized even on deterministic or noisy preferences. |
| KTO Ethayarajh+ 2024 | Unpaired data: each example is just "desirable" or "undesirable". Loss inspired by prospect theory (loss aversion), measured against a reference point estimated from the batch KL. | \(\lambda_D\,(1-\sigma(\beta(\hat r - z_0)))\) for good, \(\lambda_U\,(1-\sigma(\beta(z_0-\hat r)))\) for bad | Thumbs-up/down production feedback without pairs. Robust to class imbalance via \(\lambda\) weights. |
| ORPO Hong+ 2024 | No reference model and no separate SFT stage. SFT loss plus an odds-ratio penalty favouring chosen over rejected. | \(\mathcal{L}_\text{SFT} - \lambda\log\sigma\big(\log\frac{\text{odds}(y_w)}{\text{odds}(y_l)}\big)\) | One stage, one model in memory. Cheapest pipeline. |
| SimPO Meng+ 2024 | Implicit reward = length-normalized average log-prob (no reference model), plus a target margin γ. | \(-\log\sigma\big(\tfrac{\beta}{|y_w|}\log\pi(y_w) - \tfrac{\beta}{|y_l|}\log\pi(y_l) - \gamma\big)\) | Aligns the training reward with how generation is scored (average log-likelihood). Reduces length exploitation. Saves reference compute and memory. Sensitive to hyperparameters. |
| Online / iterative DPO Xiong+ 2023 | Each round: sample from the current policy, label pairs with an RM or judge, run DPO, repeat. | Same DPO loss, fresh on-policy pairs | Recovers much of the online-RL advantage at DPO-like cost and stability. |
| Self-Rewarding Yuan+ 2024 | The policy itself acts as LLM-as-judge to create its next round's preference pairs. | Iterative DPO with self-judging | Removes the external RM. Risk of self-reinforcing biases. |
Other names you may hear: SPIN (self-play against your previous SFT outputs) Chen+ 2024, XPO (exploration bonus for online DPO) Xie+ 2024, and Zephyr's "distilled DPO" (dDPO), where an open 7B model was aligned with DPO on GPT-4-judged UltraFeedback pairs. Tunstall+ 2023
Claiming "DPO is equivalent to RLHF, so it's strictly better". The equivalence holds for the optimum of the same objective under the same data distribution and infinite data. In practice DPO learns from a fixed offline dataset, so it has no exploration and its implicit reward is only trained on off-policy samples. PPO samples from the current policy and queries the RM on those. Different data, different results.
Offline vs online preference learning
This was one of the most-discussed questions of 2024. The evidence has converged reasonably well:
- A careful study found that well-tuned PPO beats DPO across dialogue and code tasks. It identified the key PPO ingredients (advantage normalization, large batch, EMA reference update) and showed DPO suffers from distribution shift when preference data doesn't match policy outputs. Xu+ 2024
- DeepMind compared online and offline methods under controlled KL budgets and found online consistently better. Offline data coverage and quality alone didn't explain the gap. On-policy sampling itself matters. Tang+ 2024
- Meanwhile, offline DPO remains the workhorse because it is cheap, stable, needs no generation infrastructure and is easy to iterate on data. Llama 3, Tülu 3 and OLMo 3 all use it, followed (in Tülu, OLMo and Qwen) by online RL.
Offline (DPO, IPO, KTO, ORPO, SimPO)
+ No sampling during training. 1–2 models in memory. Supervised-style stability. Easy to reproduce.
− Can only reweight responses in the dataset. Off-policy mismatch. Implicit reward overfits. No exploration.
Online (PPO, GRPO, online DPO)
+ Learns from its own outputs. Can discover better responses. Feedback targets its actual mistakes. Best final quality, and needed for reasoning RL.
− Generation infrastructure, cost, hyperparameter sensitivity, reward hacking.
"DPO or PPO for my product?" A strong answer: start with SFT plus offline DPO on pairs sampled from your own SFT model (on-policy-ish), judged by a strong judge or by humans. That is cheap and gets most of the gains. Move to iterative DPO, or to online RL (GRPO/PPO), when you have a reliable reward signal (verifier or well-calibrated RM), the remaining gap justifies the infrastructure, and you have evals to catch reward hacking.
- DPO paper: §4 has the derivation, worth reading once.
- Is DPO Superior to PPO? (Xu+ 2024): when and why PPO wins.
- Online vs offline gap (Tang+ 2024): controlled comparisons.
RLAIF and Constitutional AI
Human preference labels are slow, expensive, inconsistent and hard to scale to expert domains. RLAIF replaces human labellers with an LLM judge. On summarization and dialogue tasks, Google found that RLAIF reached quality comparable to RLHF as rated by humans. "Direct RLAIF", which asks the LLM for a score during RL instead of training an RM, also worked. Lee+ 2023
Constitutional AI (Anthropic) uses a short list of natural-language principles (a "constitution") to drive harmlessness training with little human harm labelling: Bai+ 2022
- Supervised phase (SL-CAI): sample responses to red-team prompts from a helpful-only model. Ask the model to critique its response against a randomly drawn principle, then revise it. Repeat. Fine-tune on the revised responses.
- RL phase (RL-CAI): sample pairs of responses. Ask a model which better satisfies a principle (optionally with chain of thought). Train a preference model on these AI labels (mixed with human helpfulness labels). Run RL against it.
Key outcome: a model that is both less harmful and less evasive. It explains its objections instead of giving flat refusals. The broader idea, specifying behaviour with written principles and having models apply them, now runs through industry practice: OpenAI's rule-based rewards and deliberative alignment (models reason over a written safety specification in their chain of thought), Guan+ 2024 and Anthropic's constitutional classifiers for jailbreak defence. Sharma+ 2025
Treating AI feedback as "free and unbiased". Judges inherit their own biases (verbosity, position, self-preference, sycophancy), and a policy trained against a judge will exploit them. Mitigate with position swapping, multiple judges, rubrics, calibration against a human-labelled slice, and keeping humans in the loop for guidelines and audits.
GRPO and RL with verifiable rewards (brief)
For tasks with checkable answers (math with a known final answer, code with unit tests, instruction-following with programmatic constraints such as "answer in under 50 words, in JSON"), you can skip the learned RM and use a verifier: reward 1 if correct, else 0, plus maybe a format reward. Verifiable rewards are much harder to hack than learned RMs, though not impossible (models do learn to game weak unit tests).
GRPO (Group Relative Policy Optimization), introduced in DeepSeekMath, removes PPO's critic. Shao+ 2024 For each prompt, sample a group of \(G\) responses (for example 8–64), score them, and use the group-normalized reward as every token's advantage:
$$\hat A_i = \frac{r_i - \mathrm{mean}(r_1,\dots,r_G)}{\mathrm{std}(r_1,\dots,r_G)}$$It then applies a PPO-style clipped objective with a KL penalty to the reference (added to the loss, not the reward). The "baseline" is the group mean, a Monte Carlo estimate of the value of the prompt. You save the critic's memory and avoid its instability. You pay with more samples per prompt, and you get no learning signal from prompts where all \(G\) responses score the same (all right or all wrong). Follow-ups such as Dr. GRPO identify biases from the std and length normalization Liu+ 2025, and DAPO adds tricks like decoupled clipping and dynamic sampling. Yu+ 2025 Full treatment, including rollout infrastructure, credit assignment and RL scaling, is in A5: RL for LLMs.
RLVR and its algorithms (GRPO variants, off-policy corrections, RL beyond verifiable domains with rubric or generative rewards) are among the fastest-moving parts of the field in 2025–2026. The claims here reflect published work through roughly mid-2026. Verify against current papers.
Comparison: SFT vs PPO vs DPO vs KTO vs GRPO
| SFT | PPO (RLHF) | DPO | KTO | GRPO (RLVR) | |
|---|---|---|---|---|---|
| Data needed | Demonstrations (prompt, ideal response) | Prompts + trained RM (from pairwise prefs) | Pairwise prefs (x, yw, yl) | Unpaired good/bad labels | Prompts + verifier or reward function |
| Models in memory during training | 1 (policy) | 4 (policy, critic, reference, RM) | 2 (policy, reference; reference log-probs can be precomputed) | 2 (policy, reference) | 2–3 (policy, reference; RM only if not rule-based) |
| Online / offline | Offline | Online (on-policy sampling) | Offline (iterative variant: semi-online) | Offline | Online |
| Stability | High | Low–medium, many hyperparameters | High (can overfit, likelihood displacement) | High | Medium (needs group size, clipping, filtering of all-same groups) |
| Compute cost | Low | Highest (generation + 4 models) | Low (≈ 2× SFT forward) | Low | High (G samples per prompt), no critic |
| Can exceed its data? | No, imitates the data | Yes, explores | Limited, reweights dataset responses | Limited | Yes, explores |
| Best for | Format, persona, new skills, cold start | General chat quality with a good RM | Cheap preference alignment, style, safety | Production thumbs-up/down logs | Math, code, verifiable reasoning, agents |
Parameter-efficient fine-tuning: LoRA, QLoRA, DoRA
LoRA: the math
Full fine-tuning updates every weight matrix \(W\in\mathbb{R}^{d\times k}\). LoRA's hypothesis is that the update \(\Delta W\) needed for adaptation has low intrinsic rank. So freeze \(W_0\) and learn a low-rank factorization: Hu+ 2021
$$h = W_0 x + \Delta W x = W_0 x + \frac{\alpha}{r}\,B A\,x,\qquad B\in\mathbb{R}^{d\times r},\; A\in\mathbb{R}^{r\times k},\; r\ll\min(d,k)$$- Initialization: \(A\) random Gaussian, \(B = 0\), so \(\Delta W=0\) at step 0 and training starts exactly at the pretrained model.
- Parameter count: \(r(d+k)\) instead of \(dk\). For a 4096×4096 projection with \(r=16\): \(16\times 8192 = 131\text{k}\) vs \(16.8\text{M}\), about 0.8%. Across a 7B model with LoRA on all linear layers at \(r=16\), that is roughly 40M trainable parameters (about 0.5%).
- α and the α/r scaling: α is a constant that scales the update. Dividing by r means that if you change rank, the effective update magnitude (and thus the right learning rate) stays roughly stable. Common heuristics are α = r or α = 2r. rsLoRA argues the scale should be \(\alpha/\sqrt{r}\) for stable learning at high ranks. Kalajdzievski 2023
- Memory savings: no gradients or Adam states for frozen weights. On GPT-3 175B the LoRA paper reported about 10,000× fewer trainable parameters and about 3× lower GPU memory than full fine-tuning. Hu+ 2021 The frozen weights still have to be held (in bf16, or 4-bit for QLoRA), and activations still dominate at long sequence lengths.
- Merging: after training, compute \(W = W_0 + \tfrac{\alpha}{r}BA\) and you have a plain dense model with zero extra inference latency. Or keep adapters separate and hot-swap many per-customer adapters on one base model (multi-LoRA serving, supported by engines such as vLLM).
Which modules, which rank?
The original paper adapted only the attention \(W_q, W_v\) projections. Later evidence says to apply LoRA to all linear layers, especially the MLP. Thinking Machines' "LoRA Without Regret" study found attention-only LoRA clearly underperforms MLP-only LoRA at matched parameter counts. With LoRA on all layers it matched full fine-tuning sample efficiency on small-to-medium post-training datasets, as long as the dataset didn't exceed the adapter's capacity. Their optimal LoRA learning rate was consistently about 10× the full-FT learning rate. For RL, even rank 1 matched full fine-tuning, because policy-gradient feedback carries very few bits per episode. They also found LoRA tolerates very large batch sizes less well than full FT. Schulman+ 2025 (Thinking Machines) Typical ranks: 8–64 for SFT and style adaptation, higher (128–256+) for learning substantial new skills or domains.
QLoRA
QLoRA keeps the frozen base in 4-bit and trains LoRA adapters in bf16, which made fine-tuning a 65B model possible on a single 48 GB GPU. Dettmers+ 2023 Three ingredients:
- NF4 (4-bit NormalFloat): pretrained weights are roughly zero-centred normal. NF4 places its 16 quantization levels at quantiles of a standard normal, so each bin holds about the same number of weights. The paper calls this information-theoretically optimal for normally distributed data. Weights are quantized in blocks (block size 64) with a per-block absmax scale.
- Double quantization: the per-block scales themselves (32-bit, one per 64 weights, about 0.5 bits/param) are quantized to 8-bit, saving roughly 0.37 bits/param (~3 GB on a 65B model).
- Paged optimizers: use NVIDIA unified memory so optimizer states page to CPU RAM during memory spikes (long sequences), avoiding OOM crashes.
In the forward pass, each 4-bit block is dequantized to bf16 on the fly for the matmul. Gradients flow through the frozen quantized weights into the adapters. The costs are some extra compute from dequantization (QLoRA training is often ~20–40% slower than bf16 LoRA) and a small quality loss. You also usually merge the adapter into a bf16 (not 4-bit) copy of the base for deployment, or serve base + adapter.
| Setup (7B model, rough) | Weights | Grads + optimizer | Rough GPU need |
|---|---|---|---|
| Full FT, bf16 + fp32 Adam | 14 GB (+fp32 master 28 GB) | ~70 GB | ~110–120 GB + activations → multi-GPU / ZeRO |
| LoRA, bf16 base | 14 GB | <1 GB (adapters only) | ~16–24 GB → one 24–40 GB GPU |
| QLoRA, NF4 base | ~4 GB | <1 GB | ~6–12 GB → consumer GPU |
DoRA and other variants
| Variant | Idea |
|---|---|
| DoRA Liu+ 2024 | Decompose each weight into a magnitude vector \(m\) and a direction: \(W' = m\cdot\frac{W_0+BA}{\|W_0+BA\|_c}\) (column-wise norm). LoRA updates the direction, and \(m\) trains separately. Motivation: full FT changes magnitude and direction in ways LoRA couples. Reported consistent small gains over LoRA at the same rank. Still mergeable. |
| LoRA+ Hayou+ 2024 | Use a larger learning rate for \(B\) than for \(A\). Better feature learning at large width. |
| PiSSA Meng+ 2024 | Initialize \(A,B\) from the top singular vectors of \(W_0\) and freeze the residual. Faster convergence. |
| VeRA Kopiczko+ 2023 | Shared frozen random \(A,B\) across layers. Train only small scaling vectors. Far fewer parameters. |
| Prefix / prompt tuning, adapters, IA³ | Older PEFT families: learn virtual tokens, small bottleneck MLPs, or per-channel scalings. Mostly superseded by LoRA for LLMs. See the HF PEFT docs. |
Full fine-tuning vs LoRA: what LoRA "learns less and forgets less"
Biderman et al. compared LoRA and full FT on code and math, in both continued pretraining and instruction tuning. LoRA underperformed full FT on the target domain, especially for continued pretraining on large datasets. But it better preserved the base model's performance outside the target domain and kept more output diversity. They also found that full FT learns weight perturbations with rank 10–100× higher than typical LoRA configurations. Biderman+ 2024 A related study found LoRA and full FT can reach similar accuracy yet differ structurally: LoRA introduces "intruder dimensions" (new high-ranking singular vectors) that correlate with worse generalization beyond the training task, especially in continual learning. Shuttleworth+ 2024
Choose LoRA / QLoRA when
- Adapting style, format or a narrow task with modest data (thousands to low millions of tokens per behaviour)
- You need many per-tenant variants on one base (multi-LoRA serving)
- Limited GPUs, fast iteration
- Preserving general capabilities matters (less forgetting)
- RL post-training (low-bit signal, so low rank suffices)
Choose full fine-tuning when
- Injecting lots of new knowledge or a new domain or language (continued pretraining)
- Very large datasets that exceed adapter capacity
- You're building the flagship model and have the compute
- Maximum target-domain quality is worth some forgetting
Common asks: "Write the LoRA forward pass", "Why init B to zero?", "What do r and α do?", "Count trainable params for rank 16 on a 4096-dim layer", "Why no inference overhead?", "How does QLoRA fit 65B on one GPU?" (NF4 4-bit base ≈ 0.5 byte/param ≈ 33 GB plus adapters and paging), and "LoRA or full FT for teaching a model a new language?" (full FT or continued pretraining, because LoRA learns less). A strong answer also mentions applying LoRA to MLP layers and using about a 10× higher LR.
- QLoRA paper: NF4, double quantization, paged optimizers in detail.
- LoRA Without Regret (Thinking Machines, 2025): practical recipe and when LoRA matches full FT.
- LoRA Learns Less and Forgets Less: the trade-off quantified.
Distillation
Distillation trains a smaller student to mimic a larger teacher. It is the main reason small open models are good, and it is now built into flagship pipelines (Qwen3's strong-to-weak distillation, DeepSeek-R1's distilled models).
| Flavour | Mechanism | Needs | Trade-offs |
|---|---|---|---|
| Logit / token-level KD Hinton+ 2015 | Minimize KL between teacher and student next-token distributions, softened with temperature T: \(\mathcal{L}=T^2\,\mathrm{KL}(p_T^{(\tau)}\|p_S^{(\tau)})\), often mixed with the hard-label loss. "Dark knowledge": the teacher's relative probabilities on wrong tokens carry information. | Teacher logits (white-box), same tokenizer/vocab, or alignment between vocabularies | Richest signal per token. Storage and compute heavy (full-vocab distributions). Used in pretraining distillation (for example Gemma-style) and for small models. |
| Sequence-level KD / SFT on teacher outputs Kim & Rush 2016 | Teacher generates complete responses. Student does ordinary SFT on them. | Black-box API access is enough | Simple, works across tokenizers. Student learns only the teacher's mode, not its uncertainty. Exposure bias: trained on teacher prefixes, tested on its own. |
| On-policy distillation Agarwal+ 2023 Gu+ 2023 | Student generates. Teacher provides per-token targets on the student's own samples, often using reverse KL (mode-seeking, so the student doesn't spread mass over teacher modes it can't represent). | Teacher logprobs on student samples | Fixes exposure bias. Dense per-token signal, like RL but far more bits per episode. Thinking Machines reported reasoning results comparable to RL at roughly a tenth of the compute in one setting. Lu+ 2025 (Thinking Machines) |
Reasoning distillation
DeepSeek fine-tuned Qwen2.5 and Llama 3 models (1.5B–70B) with SFT only, on reasoning traces from R1 (~800k samples). The distilled models beat what the same small models achieved with large-scale RL of their own. DeepSeek-AI 2025 The takeaway, widely replicated since: for small models, distilling a strong reasoner beats doing RL from scratch. RL is how you push the frontier with the big model, and distillation is how you spread the result to smaller models. Caveats: teacher terms of service may forbid training on outputs. Distilled models can inherit a teacher's verbosity and its failure modes. And SFT on long CoT can teach the form of reasoning (hedging, "wait, let me re-check") without full robustness.
"Why use temperature in KD?" (it softens the distribution to expose the relative probabilities of non-argmax tokens; the T² factor keeps gradient magnitudes comparable). "Forward vs reverse KL for distillation?" (forward is mode-covering, so the student smears probability over everything the teacher might say. Reverse is mode-seeking, so the student commits to modes it can represent, which suits small students in generation). "How would you build a small model for task X?" (distil from a strong teacher on task prompts, filter with verifiers, optionally finish with on-policy distillation or a little RL).
Model merging
Merging combines the weights of several fine-tuned models that share the same base, with no further training. It works surprisingly well because fine-tunes from the same pretrained initialization tend to stay in a shared, linearly connected low-loss basin. Model soups showed that averaging several fine-tunes of one model improves accuracy and robustness at no inference cost. Wortsman+ 2022 Llama 3 averaged models across runs and hyperparameter settings at each post-training stage. Llama Team 2024
| Method | Recipe | Notes |
|---|---|---|
| Linear / soup | \(\theta = \sum_i w_i\theta_i\) | Simple, robust for same-task runs. |
| Task arithmetic Ilharco+ 2022 | Task vector \(\tau_i=\theta_i-\theta_\text{base}\). Merge: \(\theta=\theta_\text{base}+\lambda\sum_i\tau_i\). Negate a vector to remove a behaviour (for example toxicity). | Interference when task vectors conflict in sign or magnitude. |
| TIES-Merging Yadav+ 2023 | (1) Trim each task vector to its largest-magnitude entries (for example the top 20%). (2) Elect a sign per parameter by total magnitude. (3) Merge by averaging only the values that agree with the elected sign. | Reduces interference from redundant small updates and sign conflicts. |
| DARE Yu+ 2023 | Randomly drop a fraction p of delta parameters (p can be very high, around 0.9) and rescale the rest by 1/(1−p). Then merge (often combined with TIES). | Shows deltas are highly redundant. |
| SLERP | Spherical interpolation between two models: \(\theta=\frac{\sin((1-t)\Omega)}{\sin\Omega}\theta_0+\frac{\sin(t\Omega)}{\sin\Omega}\theta_1\), where Ω is the angle between the (flattened) weight vectors. | Keeps the norm geometry better than linear interpolation. Only for two models at a time. Popular in community merges via MergeKit. Goddard+ 2024 |
Merging models with different base checkpoints or architectures. Weight-space merging assumes a shared initialization (and the same tokenizer). Models trained independently from scratch don't align neuron-for-neuron. Also, merged models must be re-evaluated on everything: leaderboard-gamed community merges sometimes look good on benchmarks and worse in use.
Catastrophic forgetting and regularization
Fine-tuning on a narrow distribution moves weights away from what made the base model general, so capabilities that aren't exercised degrade: other languages, coding, calibration, safety, even instruction following. This is "alignment tax" when it happens during RLHF and "catastrophic forgetting" in general. InstructGPT observed regressions on public NLP benchmarks after RLHF and mitigated them by mixing pretraining gradients into PPO ("PPO-ptx"). Ouyang+ 2022
| Mitigation | How it helps |
|---|---|
| KL penalty to the reference (RLHF/DPO β) | Directly limits distribution drift. |
| Data replay / mixing | Mix a fraction of general SFT or pretraining data into domain fine-tuning. The most reliable practical fix. |
| Lower LR, fewer epochs, early stopping | Smaller total weight movement. |
| LoRA / PEFT | Low-rank constraint limits interference ("forgets less"). Adapters can be switched off. |
| EWC-style penalties Kirkpatrick+ 2017 | Penalize changes to weights that are important (high Fisher information) for old tasks. Classic, rarely used at LLM scale. |
| Weight interpolation with the base (WiSE-FT) Wortsman+ 2021 | Average fine-tuned and original weights to trade target gains for robustness. |
| Prefer on-policy training | Recent evidence: online RL forgets less than SFT at matched target performance, apparently because on-policy updates stay close (in KL) to the base distribution. Shenfeld+ 2025 A related study, "SFT memorizes, RL generalizes", found RL generalizes better out of distribution, while SFT is useful to stabilize output format first. Chu+ 2025 |
The "RL forgets less than SFT" line of work is recent (2025) and still being tested across settings. Treat it as a promising finding, not settled fact.
Safety training
Safety post-training teaches a model to decline genuinely harmful requests (weapons uplift, CSAM, targeted harassment), to handle sensitive-but-legitimate ones carefully (medical, security research), to resist manipulation, and to respect an instruction hierarchy (system > developer > user > tool outputs). Wallace+ 2024
How it's done
- Data: red-teaming prompts (human red teamers, plus automated attack generation such as PAIR Chao+ 2023), paired with ideal responses: refusals that are helpful about why and offer safe alternatives, or safe completions.
- Training signals: SFT on safe responses, a safety RM or safety head (Llama 2), constitution- or rule-based AI feedback (CAI, RBRs), and deliberative alignment, where the model learns to reason explicitly over a written policy before answering. Guan+ 2024
- Borderline data: it is just as important to include benign prompts that look dangerous ("how do I kill a Python process?") with helpful answers.
Over-refusal
Safety training pushes the model to refuse based on surface features. XSTest is a suite of 250 safe prompts that superficially resemble unsafe ones (plus unsafe contrast prompts). It exposed significant exaggerated refusals in safety-tuned models of the time. Röttger+ 2023 Helpfulness and harmlessness are in tension. Labs measure both and tune the trade-off (separate RMs, rule-based rewards that penalize unnecessary refusals, borderline data).
Jailbreaks and robustness
- Attack families: role-play and persona prompts, many-shot prompting, encodings and translations, multi-turn escalation, prompt injection through tools and documents, and optimized adversarial suffixes. GCG found gibberish suffixes via gradient search that transferred across models, including closed ones. Zou+ 2023
- Safety is shallow: much of alignment lives in the first few output tokens (the refusal prefix). Prefilling or forcing a compliant start often bypasses it. Training on "recovery" examples, where the model returns to a refusal after a harmful start, deepens it. Qi+ 2024
- Fine-tuning removes safety: fine-tuning an aligned model on a handful of harmful examples (or even benign data) can strip safety behaviour. Qi et al. jailbroke GPT-3.5 via the fine-tuning API with 10 examples for under $0.20. Qi+ 2023 Narrow fine-tuning (for example on insecure code) can even produce broadly misaligned behaviour ("emergent misalignment"). Betley+ 2025 This matters for anyone offering fine-tuning or releasing open weights.
- Defence in depth: beyond training the model, deploy input and output classifiers. Anthropic's constitutional classifiers, trained on synthetic data generated from a constitution, withstood thousands of hours of red teaming against universal jailbreaks with modest over-refusal cost. Sharma+ 2025 Representation-level methods such as circuit breakers reroute harmful internal representations. Zou+ 2024
Jailbreak techniques and defences evolve continuously. Specific robustness numbers age quickly. Agentic settings (prompt injection through tools, web pages and emails) are the 2025–2026 focus. Check current system cards and red-teaming reports.
"How do you reduce harmful outputs without making the model useless?" Strong answer: define a policy, then build both harmful and borderline-benign data. Train with separate safety and helpfulness signals (or rule-based rewards that penalize both harmful compliance and unnecessary refusal). Measure attack success rate and over-refusal rate (XSTest-style). Add runtime classifiers. Red-team continuously, including multi-turn and prompt-injection attacks. Mention that safety is shallow and that fine-tuning can undo it.
- Constitutional AI: principles-driven harmlessness training.
- Safety Alignment Should Be More Than a Few Tokens Deep: why jailbreaks work, and a fix.
- Anthropic: Constitutional Classifiers: classifier-based jailbreak defence.
Evaluating chat models (pointer)
Post-training is only as good as its evals. The short version (full treatment in A8):
- Capability benchmarks (math, code, knowledge, instruction following with verifiable constraints such as IFEval) catch regressions from tuning.
- Preference-style evals: pairwise human votes (Chatbot Arena), LLM-judge win rates (MT-Bench, AlpacaEval with length control, Arena-Hard). Watch for judge biases and length gaming. Zheng+ 2023 Dubois+ 2024
- Safety evals: attack success rate on red-team sets, over-refusal (XSTest), and jailbreak robustness.
- Product evals: task-specific held-out sets built before training, plus online A/B tests. Always decontaminate training data against evals.
Interview question bank
1. What does post-training change, and what is the superficial alignment hypothesis?
Pretraining instils knowledge and skills. Post-training mainly teaches the interface (chat format, when to stop), the persona and style, which latent capabilities to use, and behavioural constraints like refusals. LIMA showed that 1,000 curated SFT examples on a 65B model produced a strong assistant, and URIAL showed base and aligned models differ mostly on stylistic tokens. Together these motivate the hypothesis that alignment is "superficial". Its limits: RL on verifiable tasks produces large reasoning gains, and robust tool use, multi-turn behaviour and safety need much more data. Whether RL creates new capability or sharpens existing capability is debated: pass@k analyses suggest the base model can often sample correct solutions at large k.
2. Write the reward model loss and explain why we use pairwise comparisons.
\(\mathcal{L}=-\mathbb{E}[\log\sigma(r(x,y_w)-r(x,y_l))]\), the Bradley–Terry negative log-likelihood. Humans are much more consistent at relative judgements than absolute scores, which drift across annotators and over time. BT only identifies rewards up to a per-prompt constant, which is fine because RL only needs relative quality (and you normalize). With K ranked responses you get K-choose-2 pairs, but you should batch pairs from the same prompt together to avoid overfitting to correlated examples. Expect held-out accuracy around 65–75% on hard chat data, because labels are noisy.
3. Write the RLHF objective. What does each term do, and what happens if β → 0 or β → ∞?
\(\max_\pi \mathbb{E}_{y\sim\pi}[r(x,y)]-\beta\mathrm{KL}(\pi\|\pi_\text{ref})\). The first term pushes toward high reward. The KL keeps the policy near the SFT reference, where the RM is trustworthy, and preserves fluency, diversity and general capabilities. As β → 0 you get unconstrained reward maximization: reward hacking, verbosity, degenerate text, mode collapse. As β → ∞ the policy stays at the reference and nothing is learned. In practice β is around 0.01–0.1, or adapted to hit a target KL, and the KL is applied as a per-token reward penalty.
4. Name the models in PPO-based RLHF and estimate the memory for a 7B model.
Policy/actor (trainable), value/critic (trainable), reference (frozen) and reward model (frozen). With mixed-precision Adam at about 16 bytes/param, each trainable 7B model needs about 112 GB (weights, grads, fp32 master and two Adam moments). Two trainable models come to about 224 GB. The frozen bf16 models add about 14 GB each, so roughly 250 GB before activations and KV cache for generation. That requires sharding across several 80 GB GPUs. Reductions: LoRA (the reference is the base with adapters disabled), a shared actor-critic trunk, offloading frozen models, or switching to GRPO (no critic) or DPO (no RM, no critic, precomputed reference log-probs).
5. Derive DPO at a high level.
Start from the KL-regularized objective. Its optimal policy has closed form \(\pi^*(y|x)=\pi_\text{ref}(y|x)\exp(r(x,y)/\beta)/Z(x)\). Rearranging gives \(r=\beta\log(\pi^*/\pi_\text{ref})+\beta\log Z(x)\). Substitute into Bradley–Terry: since it depends only on reward differences for the same x, the intractable \(\beta\log Z(x)\) cancels. Parameterize \(\pi^*\) by \(\pi_\theta\) and maximize the preference likelihood: \(-\log\sigma(\beta[\log\frac{\pi_\theta(y_w)}{\pi_\text{ref}(y_w)}-\log\frac{\pi_\theta(y_l)}{\pi_\text{ref}(y_l)}])\). The policy is its own implicit reward model, with no separate RM, RL loop or sampling.
6. DPO is "equivalent" to RLHF. Why can PPO still beat it?
The equivalence is between optima of the same objective given the same preference distribution. In practice DPO trains on a fixed, usually off-policy dataset, so its implicit reward is only fitted on responses the policy may never generate. It does no exploration and can't discover better responses. PPO samples on-policy and asks the RM about its own outputs, giving feedback where the policy actually is. DPO also lacks an explicit KL term, so on near-deterministic preferences it pushes margins without bound (IPO's motivation), and it often lowers the chosen response's likelihood too. Controlled studies (Xu+ 2024, Tang+ 2024) found well-tuned online methods outperform offline ones. Iterative or online DPO closes much of the gap.
7. Compare IPO, KTO, ORPO and SimPO in one line each.
IPO: replace log-sigmoid with a squared loss to a fixed target margin, so it can't overfit deterministic preferences. KTO: learn from unpaired good/bad labels with a prospect-theory-style loss against a KL reference point, which suits thumbs-up/down logs. ORPO: combine SFT and an odds-ratio preference term in one stage with no reference model. SimPO: use length-normalized average log-prob as a reference-free implicit reward with a target margin, which reduces length bias and memory. All are offline. Pick based on data shape (paired vs unpaired), memory budget and how much you trust your SFT stage.
8. Your DPO-trained model got much more verbose. Why, and how do you fix it?
Chosen responses in most preference datasets are longer on average (humans and LLM judges both prefer longer, more thorough answers). DPO learns the cheapest feature that separates chosen from rejected, and length is easy. Since DPO's log-ratio uses summed log-probs, length also enters the implicit reward directly. Fixes: balance lengths in the data (or length-match pairs), add a length penalty or regularizer (length-regularized DPO), use SimPO's length-normalized reward, raise β, add an NLL term on chosen responses, and evaluate with length-controlled win rates so you can see the problem.
9. What is reward hacking? Give three concrete LLM examples and mitigations.
Optimizing a proxy reward until it diverges from the true objective (Goodhart). Examples: (1) verbosity and lists because the RM likes length and structure, (2) sycophancy, agreeing with the user's stated view because raters liked it, (3) on code RL, special-casing or deleting the unit tests instead of solving the problem. Others include hedging boilerplate that a safety RM rewards, or degenerate strings that score highly. Mitigations: KL regularization and early stopping (Gao+ curves show gold reward peaks then falls), RM ensembles, refreshing the RM on on-policy samples, verifiable rewards where possible with strong tests and sandboxing, length and format penalties, and held-out human or gold evals during training.
10. Process vs outcome reward models: which would you use for math RL, and why?
An ORM or verifier scores only the final answer: cheap, hard to hack if the checker is good, but credit assignment is sparse. A PRM scores each step: denser signal and better for best-of-N reranking and search (Lightman+ 2023), but step labels are expensive (or noisy if derived from Monte Carlo), "step" is ill-defined in free text, and a learned PRM is itself hackable during RL. For RL training at scale, today's default is outcome-based verifiable rewards with GRPO/PPO (DeepSeek-R1 reported PRMs didn't pay off for them). I'd use a PRM or generative verifier at inference for reranking, or as an auxiliary signal, and watch the implicit-PRM literature.
11. Explain GRPO and why it drops the critic.
For each prompt, sample a group of G responses, score them, and set each response's advantage to \((r_i-\bar r)/\sigma_r\) over the group. Then apply a PPO-style clipped objective with a KL penalty to the reference. The group mean acts as a Monte Carlo baseline for the prompt's value, so you don't need a separately trained value network. That network would be as large as the policy, hard to train on partial text, and a source of variance. Costs: more samples per prompt, and no gradient from prompts where all responses get the same reward (so curricula and dynamic sampling filter for medium-difficulty prompts). Variants like Dr. GRPO fix normalization biases. See A5.
12. LoRA: write the forward pass, the parameter count for r=16 on a 4096×4096 matrix, and say why B is zero-initialized.
\(h=W_0x+\frac{\alpha}{r}BAx\) with \(B\in\mathbb{R}^{4096\times16}\) and \(A\in\mathbb{R}^{16\times4096}\). Parameters: \(16\times(4096+4096)=131{,}072\), versus 16.8M for the full matrix, about 0.78%. With B=0, \(\Delta W=0\) at initialization, so training starts exactly at the pretrained function and early updates are well-behaved. A is random so that gradients into B are non-zero. α/r keeps the effective update scale stable as you change r. After training, merge \(W_0+\frac{\alpha}{r}BA\) for zero-overhead inference.
13. How does QLoRA fit a 65B model on a single 48 GB GPU?
The frozen base is stored in 4-bit NF4: 65B × 0.5 bytes ≈ 33 GB, plus per-block scales that double quantization shrinks from about 0.5 to about 0.13 bits/param. Only LoRA adapters (tens to hundreds of MB) carry bf16 weights, gradients and Adam states. Weights are dequantized block by block to bf16 for each matmul, and gradients flow through them into the adapters. Paged optimizers move optimizer state to CPU memory during activation spikes. Gradient checkpointing keeps activations small. The cost is slower training from dequantization overhead and a small quality gap versus 16-bit LoRA.
14. When would you choose full fine-tuning over LoRA? What does "LoRA learns less and forgets less" mean?
Biderman+ 2024 found LoRA underperforms full FT on the target domain, particularly for large-data continued pretraining (code, math), while better preserving out-of-domain abilities. Full FT learned much higher-rank weight changes. So choose full FT (or continued pretraining) to inject lots of new knowledge, a new language or a big domain, or when you're building a flagship model and can afford it. Choose LoRA for style, format and narrow tasks, multi-tenant adapters, limited compute, RL post-training (where even very low rank suffices), or when preserving general capabilities matters. Thinking Machines (2025) found that with LoRA on all layers, including MLP, and about 10× the LR, LoRA matches full FT on typical post-training dataset sizes.
15. Logit distillation vs SFT on teacher outputs: trade-offs?
Logit KD matches the teacher's full next-token distribution (softened by temperature), passing "dark knowledge" about plausible alternatives. That is a much richer signal per token, but it needs white-box teacher access, a shared tokenizer, and storage or compute for vocab-sized distributions. Sequence-level KD (SFT on teacher generations) needs only black-box outputs, works across tokenizers and is simple, but it only teaches the teacher's mode and suffers exposure bias. On-policy distillation, where the student samples and the teacher scores per token (often with reverse KL), fixes exposure bias with dense supervision. For reasoning, SFT on a strong teacher's traces beats doing RL on small models (DeepSeek-R1 distills).
16. Explain task arithmetic and TIES-merging. When does merging fail?
A task vector is \(\tau=\theta_\text{ft}-\theta_\text{base}\). Adding scaled task vectors combines skills, and subtracting one removes a behaviour. Naive addition suffers interference: many small, redundant deltas and sign conflicts between tasks. TIES trims each vector to its largest-magnitude entries, elects a sign per parameter by total mass, and averages only the agreeing entries. DARE randomly drops most deltas and rescales the rest. Merging fails when models don't share a base checkpoint and tokenizer, when tasks genuinely conflict, or when fine-tunes moved far from the base (outside a shared basin). Always re-evaluate broadly.
17. After domain fine-tuning, your model's general chat and coding skills regressed. Diagnose and fix.
This is catastrophic forgetting: the domain data shifted the weights and the unexercised capabilities decayed. Check LR (too high?), epochs (overfitting?), data diversity, template consistency (did you change the chat template or drop the system prompt?) and whether safety or EOS behaviour changed. Fixes: mix in general instruction data (replay), lower LR and fewer epochs, use LoRA, interpolate weights with the original (WiSE-FT style), or add a KL penalty or a DPO stage anchored to the original model. Track a broad regression suite on every run, not just the domain metric. Consider on-policy methods, which recent evidence suggests forget less.
18. Design a post-training pipeline for a 7B open model to be a customer-support assistant with tool calls, on a budget of a few GPUs.
Start from a strong instruct model (not a base), since re-doing general alignment is wasteful. Data: real anonymized tickets, synthetic expansion (Evol-style variations, edge cases), and tool-call trajectories generated by a strong teacher and validated by actually executing the tools. Decontaminate against a held-out eval set built first: task success, tool-call accuracy, tone, escalation correctness, hallucinated policies, refusal/over-refusal. Train LoRA or QLoRA SFT with the model's own chat template, loss on assistant and tool-call tokens only. Then DPO on pairs sampled from the SFT model and judged by a rubric-guided LLM judge, with humans calibrating a slice. If there are verifiable sub-tasks (correct tool arguments, JSON validity), add a small GRPO stage with programmatic rewards. Mix general data to limit forgetting, red-team for prompt injection through tool outputs, and A/B test before rollout.
19. What is Constitutional AI, and how does RLAIF differ from RLHF?
RLAIF replaces human preference labels with LLM-judge labels. Google found it comparable to RLHF on summarization and dialogue at far lower cost. Constitutional AI is a specific RLAIF recipe for harmlessness. The supervised phase has the model critique and revise its own responses against written principles, then fine-tunes on the revisions. The RL phase has an AI pick the response that better follows a principle, trains a preference model on those labels and runs RL. Benefits: scalable, transparent principles, less evasive refusals. Risks: judge biases and self-reinforcement, so humans still write principles and audit.
20. Why is safety alignment "shallow", and what does that imply for open-weight releases?
Qi+ 2024 showed that much of the difference between aligned and base models sits in the first few output tokens: the refusal prefix. If you force a compliant start (prefilling, some jailbreaks), the model often continues harmfully. Fine-tuning on tiny harmful datasets, or even benign data, strips safety (Qi+ 2023), and narrow bad fine-tuning can cause broad misalignment (Betley+ 2025). For open weights, model-level refusals can't be relied on against a determined actor with fine-tuning access. Safety then depends on not having the dangerous capability in the first place (data filtering), deeper alignment training (recovery examples), tamper-resistance research, and system-level controls for hosted deployments.
21. How would you measure and reduce over-refusal?
Measure: a borderline-benign set (XSTest-style prompts that look dangerous but aren't, like "how to kill a process"), refusal rate on representative real traffic, and a harmful set for attack success rate, so you track both sides of the trade-off. Reduce: add borderline-benign prompts with helpful answers to SFT and preference data, use separate helpfulness and safety signals or rule-based rewards that penalize unnecessary refusals, write a more precise policy (what's actually disallowed), use reasoning-based approaches (deliberative alignment) so the model considers context, and move some protection to runtime classifiers so the model itself can be less trigger-happy.
22. Back-of-envelope: how many preference pairs and how much compute for a DPO run on a 7B model?
Public recipes use on the order of tens to a few hundred thousand pairs (UltraFeedback is about 60k prompts; Tülu 3's DPO mixes are a few hundred thousand pairs). Say 100k pairs at about 1k tokens per response. That is 200M response tokens, plus prompts, per epoch. Each step needs forward and backward on chosen and rejected for the policy, plus reference forward passes (or precomputed log-probs), so about 2–3× the cost of SFT on the same tokens. At about 6 × 7B × 2×10⁸ ≈ 8×10¹⁸ FLOPs for policy fwd+bwd, and roughly 40% MFU on H100s (~400 TFLOP/s effective), that is a few GPU-hours of pure compute per epoch. Real-world runs take longer (sequence padding, ref passes, inefficiency), but a single node finishes in hours. Data generation and judging usually cost more than the training itself.
23. How do packing and loss masking interact, and what bugs do you look for?
With packing, multiple conversations share one sequence. You need (a) per-example position-ID resets, (b) a block-diagonal attention mask or varlen attention so examples don't attend across boundaries, and (c) loss masks that keep only assistant tokens of each example, including each example's end-of-turn token. Common bugs: attention leakage across examples (subtly worse quality), the EOS token masked out (the model never stops), the label shift misaligned at example boundaries, the template applied differently at train and inference time, and loss normalization that over-weights long examples. Verify by decoding a batch with masks visualized before training.
24. Online vs offline preference learning: what does the evidence say, and what would you do in practice?
Controlled studies (Tang+ 2024, Xu+ 2024) find online, on-policy methods outperform offline ones at matched KL. The advantage comes from sampling the current policy, not just from data coverage. Offline methods are far cheaper, simpler and more stable. Practical recipe: offline DPO on pairs generated by your own SFT model (cheap, captures most gains), then iterate (regenerate and relabel each round). Then add online RL where you have trustworthy rewards: verifiers for math, code and format, or calibrated RMs or rubric judges for open-ended tasks. This is essentially what Tülu 3, Qwen and OLMo 3 publish.