LLM Post-Training - From Base Model to Assistant: SFT, Preference Tuning, PEFT, Distillation & Evals
Hub: base model -> assistant pipeline (SFT, RLHF/PPO, DPO-family and GRPO, LoRA/QLoRA, distillation, eval gate); techniques compared by signal needed, cost and failure mode; runnable pipeline toy + technique picker; decision chart.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Why is a pretrained model not yet an assistant?
Answer
It predicts internet text: it continues questions, ignores formats and has no notion of which answer people prefer.
L2
What does SFT teach?
Answer
Mostly format, persona, task skills and tool-call syntax, from (prompt, ideal answer) demonstrations.
L3
Why add preference tuning after SFT?
Answer
SFT can only imitate; many qualities are easier to judge than to write, and comparisons capture them.
L4
What do PPO, DPO and GRPO have in common?
Answer
All optimise reward minus a KL leash to a reference model.
L5
When is GRPO the right pick?
Answer
When a program can grade answers, such as unit tests, exact numeric answers or a SQL result check.
L6
Why do LoRA and QLoRA matter?
Answer
Full fine-tuning needs about 16 bytes per parameter; LoRA trains well under 1 percent of parameters and QLoRA stores the base in 4-bit NF4.
L7
Why is an eval gate part of the pipeline?
Answer
Every stage optimises a proxy and can regress something the last stage fixed.
Failure modes
Preference tuning without SFT
The base model rarely produces well-formed answers, so pairs compare two bad completions and the signal is noise.
Over-optimised reward model
RM scores keep rising after real quality peaks.
Template mismatch in serving
Different special tokens at inference silently degrade every answer.
Misconceptions
Falling training loss means the model improved.
SFT loss falls while memorising and DPO loss can fall while the chosen answer gets less likely; track evals and chosen log-probs.
Distillation is a quality play.
The student is capped by the teacher and inherits its mistakes; it is a cost play.
PPO is always better than DPO.
With a few thousand fixed pairs, DPO is usually as good and far simpler.
Interviewer traps
Jumping to fine-tuning when prompting already passes evals.
Stop at prompting plus retrieval; it is reversible and free to iterate.
Full fine-tuning a 70B model for a tone change.
Use a LoRA adapter; many adapters can share one base.
Design scenario
Same prompt for every reader.
Requirements
Consistent output format, correct queries on the test DB, no regressions in safety refusals, and a release gate.
Failure assumptions
- The reward checker can be gamed by special-casing tests.
- The serving template differs from training.
- A gain on the target benchmark hides a safety regression.
Constraints
- One 80 GB GPU for training.
- No budget for human preference labels.
Prompt
Plan post-training for a SQL assistant on an open 8B model with 20,000 demonstrations and a test database.
API
Which data formats and chat template does each stage use?
Data
Which demos, graders and regression slices exist, and how is contamination checked?
Architecture
Which stages run in what order, and where does the eval gate sit?
Overview
A pretrained LLM is a very good next-token predictor for internet text. It is not yet an assistant: it continues a question instead of answering it, ignores formats, and has no notion of which of two plausible answers people prefer. Post-training is everything done after pretraining to turn that base model into a deployable product: supervised fine-tuning (SFT) on demonstrations, preference tuning (RLHF with PPO, or DPO-family and GRPO-style methods), often with parameter-efficient fine-tuning (LoRA/QLoRA) to make it affordable, distillation to ship a smaller model, and an eval gate that decides whether the result is allowed near users. This hub gives the whole pipeline, compares the techniques by what signal they need and what they cost, and ends with a decision chart. Each sibling page goes deep on one stage.
Scope: training-side only. Serving the result (quantization, prompt caching, multi-LoRA serving, capacity) lives in the sibling cluster LLM Inference in Production (
llm-inference-production-quantization-caching-multilora-capacity), and runtime internals (KV cache, PagedAttention, vLLM/SGLang) live in the existing inference pages linked under Related.
The pipeline in one picture
Diagram 1: the post-training pipeline, with the failure path most teams hit.
Decisions
- 1
Step 1: pretrained base model (next-token predictor on web text)
- nextStep 2: SFT on curated demonstrations (chat template, loss on assistant tokens only)
- 2
Step 2: SFT on curated demonstrations (chat template, loss on assistant tokens only)
- nextStep 3: preference signal collected (human or AI comparisons, or a program that grades answers)
- 3
Step 3: preference signal collected (human or AI comparisons, or a program that grades answers)
- nextStep 4: which optimizer?
- ?
Step 4: which optimizer?
- nextStep 5a: DPO-family update against a frozen reference
- nextStep 5b: RLHF with PPO plus KL penalty
- nextStep 5c: GRPO-style RL on sampled groups
- 5
Step 5a: DPO-family update against a frozen reference
- nextStep 6: optional distillation into a smaller student
- 6
Step 5b: RLHF with PPO plus KL penalty
- nextStep 6: optional distillation into a smaller student
- 7
Step 5c: GRPO-style RL on sampled groups
- nextStep 6: optional distillation into a smaller student
- 8
Step 6: optional distillation into a smaller student
- nextStep 7: eval gate - benchmarks, judge, regression and safety slices
- ?
Step 7: eval gate - benchmarks, judge, regression and safety slices
- nextStep 8: deploy behind canary, keep monitoring
- nextFailure path: regression found (safety slice drop, reward hacking, format breakage)
- 10
Step 8: deploy behind canary, keep monitoring
- 11
Failure path: regression found (safety slice drop, reward hacking, format breakage)
- nextFix data or hyperparameters at the stage that caused it, never patch with a bigger beta blindly
- 12
Fix data or hyperparameters at the stage that caused it, never patch with a bigger beta blindly
- nextStep 2: SFT on curated demonstrations (chat template, loss on assistant tokens only)
Lesson map
LLM Post-Training - From Base Model to Assistant: SFT, Preference Tuning, PEFT, Distillation & Evals
Hub: base model -> assistant pipeline (SFT, RLHF/PPO, DPO-family and GRPO, LoRA/QLoRA, distillation, eval gate); techniques compared by signal needed, cost and failure mode; runnable pipeline toy + technique picker; decision chart.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: pretrained base model (next-token predictor on web text)"] s2["Step 2: SFT on curated demonstrations (chat template, loss on assistant tokens only)"] s3["Step 3: preference signal collected (human or AI comparisons, or a program that grades answers)"] s4["Step 4: which optimizer?"] s5a["Step 5a: DPO-family update against a frozen reference"] s5b["Step 5b: RLHF with PPO plus KL penalty"] s5c["Step 5c: GRPO-style RL on sampled groups"] s6["Step 6: optional distillation into a smaller student"] s7["Step 7: eval gate - benchmarks, judge, regression and safety slices"] s8["Step 8: deploy behind canary, keep monitoring"] f1["Failure path: regression found (safety slice drop, reward hacking, format breakage)"] f2["Fix data or hyperparameters at the stage that caused it, never patch with a bigger beta blindly"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s4 -->|continues| s5a s4 -->|continues| s5b s4 -->|continues| s5c s5a -->|continues| s6 s5b -->|continues| s6 s5c -->|continues| s6 s6 -->|continues| s7 s7 -->|continues| s8 s7 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s2
| Stage | Signal it needs | What it changes | Typical cost | Main risk |
|---|---|---|---|---|
| SFT (instruction tuning) | (prompt, ideal answer) demonstrations | Format, persona, task skills, tool-call syntax | Low to medium: one supervised pass | Imitates noise in demos; forgets other skills |
| Reward model + PPO (RLHF) | Pairwise comparisons, then online sampling | Fine-grained preferences beyond what demos show | Highest: 4 models in memory, online generation | Reward hacking, instability, infra complexity |
| DPO family (DPO, IPO, ORPO, SimPO, KTO) | Offline comparisons (KTO: thumbs up/down) | Same goal as RLHF, closed-form loss | Medium: like SFT plus a reference forward pass | Likelihood displacement, off-policy staleness |
| GRPO-style RL with verifiable rewards | A program that scores outputs (tests, exact answers) | Reasoning and correctness on checkable tasks | High sampling cost, no critic | Reward gaming of the checker, zero-signal groups |
| LoRA / QLoRA | Any of the above | Same as the method it wraps, with far fewer trainable params | Lowest memory | Under-capacity for big domain shifts |
| Distillation | A stronger teacher's outputs or logits | Moves capability into a cheaper student | Teacher inference + student training | Student inherits teacher errors, capped by teacher |
| Evals | Held-out tasks, judges, regression sets | Nothing: it decides | Ongoing | Contamination, judge bias, Goodhart |
Fixed pairs or a program that grades answers?
Prefer
DPO-family on offline pairs
Turn the RLHF objective into a supervised loss on (chosen, rejected) pairs.
- No reward model, no critic, no online sampling.
- Beta decides the drift: beta=0.1 reached KL 0.562 and collapsed to one answer, beta=1.0 kept KL 0.470.
- Limited to the pairs you have, which go stale as the policy moves.
Alternative
GRPO-style RL with verifiable rewards
Sample a group per prompt and grade each answer with a program.
- Online: keeps finding new failures to fix.
- No critic: the group mean is the baseline.
- Higher sampling cost, and the grader itself can be gamed.
From base model to shipped assistant
Diagram 1 condensed: the post-training pipeline and its failure path.
- 1
Supervised fine-tuning
Demonstrations in a chat template, loss on assistant tokens. - 2
Preference tuning
PPO-RLHF, DPO-family or GRPO, all with a KL leash to a reference. - 3
Make it affordable
LoRA or QLoRA wraps whichever method you use. - 4
Distill if serving cost matters
A smaller student learns from the teacher at a known gap. - 5
Eval gate
Capability, regression slices, safety and contamination decide the release.
Why each stage exists, compared
SFT vs prompting. If a strong prompt plus retrieval already passes your evals, stop: prompting is reversible and free to iterate. SFT wins when you need consistent format at scale, shorter prompts (cheaper per request), a capability the base model only shows inconsistently, or a smaller model to behave like a larger prompted one.
Preference tuning vs more SFT. SFT can only imitate; it learns "say this" but not "this is better than that." Many qualities are easier to judge than to write: helpfulness, tone, safe refusals, concise answers. Comparisons capture those. A second reason: SFT on a single gold answer treats every other answer as equally wrong, while preference losses push specifically away from the bad behaviour you observed.
PPO vs DPO vs GRPO. All three optimise "reward minus a KL leash to a reference model." PPO does it online with a learned reward model and a critic, so it can keep improving on fresh samples but needs four models and careful tuning. DPO rewrites the same objective into a supervised loss on fixed pairs: far simpler and cheaper, but it only learns from the pairs you have, which go stale as the policy moves. GRPO keeps online RL but drops the critic by comparing each sample to its own group; it shines when a program can grade answers (math, code, SQL against a test database), which is how recent reasoning models were trained.
Full fine-tuning vs LoRA vs QLoRA. Full fine-tuning updates every weight and needs roughly 16 bytes per parameter for weights, gradients and Adam states. LoRA freezes the base and learns a low-rank update, cutting trainable parameters to well under 1 percent; QLoRA additionally stores the frozen base in 4-bit NF4 so a 70B model can be tuned on one 80 GB GPU. Quality is usually close for behaviour and format changes; full fine-tuning still wins for large knowledge or domain shifts.
Distillation vs training a small model directly. A strong teacher produces better targets than most human datasets at scale, and its full probability distribution carries more information per example than a single label. The student is capped by the teacher and inherits its mistakes, so distillation is a cost play, not a quality play.
Evals vs vibes. Every stage optimises a proxy. Without a gate that checks capability, regression slices, safety and contamination, post-training quietly trades one quality for another.
Runnable: the whole pipeline on a four-answer toy
The "model" below is a softmax over four whole responses to one prompt, small enough to print, but SFT and DPO use the same update shapes as real trainers. Watch SFT move mass to the demonstration, DPO sharpen the chosen-vs-rejected gap, and beta decide how far the policy drifts from its reference.
"""Toy post-training pipeline on ONE prompt with FOUR candidate responses.
The "model" is just a softmax over 4 logits, one per whole response. That is far
smaller than a real LLM, but the update rules are the same shapes:
SFT : minimise -log pi(demo) (imitate a demonstration)
DPO : minimise -log sigmoid(beta * margin) (prefer chosen over rejected,
margin measured relative to a frozen reference policy)
We print probabilities and KL(pi || reference) after each stage.
Pure stdlib, deterministic.
"""
import math
RESPONSES = ["web-page continuation", "helpful + correct", "confident + wrong", "refusal"]
def softmax(z):
m = max(z)
e = [math.exp(v - m) for v in z]
s = sum(e)
return [v / s for v in e]
def kl(p, q):
"""KL(p || q) in nats: how far p has drifted from q."""
return sum(pi * math.log(pi / qi) for pi, qi in zip(p, q) if pi > 0)
def show(tag, z, ref=None):
p = softmax(z)
cells = " ".join(f"{r[:18]:>18}={v:5.3f}" for r, v in zip(RESPONSES, p))
extra = f" KL(pi||ref)={kl(p, softmax(ref)):.3f}" if ref else ""
print(f"{tag:<26}{cells}{extra}")
# Stage 0: a pretrained base model mostly continues text like the web.
base = [2.0, 0.3, 0.5, -1.0]
show("0 base (pretrained)", base)
# Stage 1: SFT on a demonstration of the helpful answer (index 1).
sft = base[:]
lr = 0.5
for _ in range(4): # a short SFT run: the demo becomes the most likely answer
p = softmax(sft)
grad = [pi - (1.0 if i == 1 else 0.0) for i, pi in enumerate(p)] # d(-log p_demo)/dz
sft = [z - lr * g for z, g in zip(sft, grad)]
show("1 after SFT", sft, ref=base)
# Stage 2: DPO with reference = frozen SFT model. Pair: helpful (1) > confident wrong (2).
def dpo(policy, ref, chosen, rejected, beta, steps, lr):
z = policy[:]
for _ in range(steps):
lp, lr_ = [math.log(v) for v in softmax(z)], [math.log(v) for v in softmax(ref)]
margin = (lp[chosen] - lr_[chosen]) - (lp[rejected] - lr_[rejected])
w = 1 / (1 + math.exp(beta * margin)) # sigmoid(-beta*margin): big when the pair is still wrong
# For a softmax over whole responses, d(margin)/dz = e_chosen - e_rejected.
z[chosen] += lr * beta * w
z[rejected] -= lr * beta * w
return z
for beta in (0.1, 1.0):
out = dpo(sft, sft, chosen=1, rejected=2, beta=beta, steps=300, lr=1.0)
lp, lq = [math.log(v) for v in softmax(out)], [math.log(v) for v in softmax(sft)]
margin = (lp[1] - lq[1]) - (lp[2] - lq[2])
show(f"2 DPO beta={beta}", out, ref=sft)
print(f"{'':26}log-ratio margin={margin:5.2f} (loss saturates once beta*margin >> 1)")
# Failure path: no anchor. Pure "maximise the reward" drifts without limit.
greedy = sft[:]
for _ in range(300):
greedy[1] += 0.1 # keep pushing the reward model's favourite, no KL term, no saturation
show("X unanchored reward push", greedy, ref=sft)
print("\nTakeaways: SFT moves mass to the demo; DPO sharpens the chosen-vs-rejected gap;")
print("small beta needs a large log-ratio margin before the loss saturates, so it drifts further from the")
print("reference and collapses to one answer; large beta saturates early and keeps some diversity.")
print("Note KL(pi||ref) caps at -log ref(y) once pi is one-hot, so watch margins and entropy too.")Output:
0 base (pretrained) web-page continuat=0.687 helpful + correct=0.126 confident + wrong=0.153 refusal=0.034
1 after SFT web-page continuat=0.272 helpful + correct=0.569 confident + wrong=0.124 refusal=0.035 KL(pi||ref)=0.583
2 DPO beta=0.1 web-page continuat=0.000 helpful + correct=1.000 confident + wrong=0.000 refusal=0.000 KL(pi||ref)=0.562
log-ratio margin=16.75 (loss saturates once beta*margin >> 1)
2 DPO beta=1.0 web-page continuat=0.019 helpful + correct=0.978 confident + wrong=0.000 refusal=0.002 KL(pi||ref)=0.470
log-ratio margin= 6.40 (loss saturates once beta*margin >> 1)
X unanchored reward push web-page continuat=0.000 helpful + correct=1.000 confident + wrong=0.000 refusal=0.000 KL(pi||ref)=0.563
Takeaways: SFT moves mass to the demo; DPO sharpens the chosen-vs-rejected gap;
small beta needs a large log-ratio margin before the loss saturates, so it drifts further from the
reference and collapses to one answer; large beta saturates early and keeps some diversity.
Note KL(pi||ref) caps at -log ref(y) once pi is one-hot, so watch margins and entropy too.Runnable: the decision chart as code
The same logic as the decision chart below, applied to four realistic projects. The memory rule of thumb (about 16 bytes per parameter for a full fine-tune with AdamW) is what pushes larger models toward LoRA and QLoRA.
// Decision chart for "which post-training technique?" written as code.
// Each rule mirrors one diamond in the hub's decision chart. Run: tsc --strict + node.
type Signals = {
promptingAndRagFail: boolean; // did a strong prompt + retrieval already solve it?
haveDemonstrations: number; // count of high-quality (prompt, ideal answer) pairs
havePreferencePairs: number; // count of (prompt, chosen, rejected) comparisons
verifiableReward: boolean; // can a program grade outputs (tests pass, exact answer)?
needsOwnBaseModel: boolean; // must ship a smaller/cheaper model of your own?
gpuMemoryGB: number; // biggest single-node memory you can use for training
modelParamsB: number; // model size in billions of parameters
};
type Plan = { steps: string[]; why: string[] };
function pick(s: Signals): Plan {
const steps: string[] = [];
const why: string[] = [];
if (!s.promptingAndRagFail) {
return { steps: ["prompting + RAG + evals"], why: ["cheapest lever works; do not train yet"] };
}
// Rough full fine-tune memory: ~16 bytes/param for weights+grads+Adam states (activations extra).
const fullFtGB = s.modelParamsB * 16;
const peft = fullFtGB > s.gpuMemoryGB;
const qlora = s.modelParamsB * 2 > s.gpuMemoryGB * 0.8; // bf16 base alone barely fits -> 4-bit base
const tuner = !peft ? "full fine-tune" : qlora ? "QLoRA (4-bit base + LoRA)" : "LoRA";
if (s.haveDemonstrations >= 1000) {
steps.push(`SFT via ${tuner}`);
why.push(`${s.haveDemonstrations} demos teach format/behaviour; full FT needs ~${fullFtGB} GB vs ${s.gpuMemoryGB} GB`);
} else {
steps.push("collect or synthesise demos first (or distill from a stronger teacher)");
why.push("SFT on a few hundred noisy demos mostly teaches noise");
}
if (s.verifiableReward) {
steps.push("RL with verifiable rewards (GRPO-style, no reward model)");
why.push("a program grader is harder to flatter than a learned reward model (still audit for test hacking)");
} else if (s.havePreferencePairs >= 5000) {
steps.push("DPO-family preference tuning (offline)");
why.push("pairs exist; no reward model or online sampling needed");
} else if (s.havePreferencePairs > 0) {
steps.push("reward model + PPO only if you can afford online RL; else gather more pairs");
why.push("few pairs: a reward model generalises them, but RLHF is the most expensive loop");
}
if (s.needsOwnBaseModel) {
steps.push("distill: teacher generates, smaller student is trained (SFT on outputs or logits)");
why.push("cheaper serving at a known quality gap");
}
steps.push("eval gate (benchmarks + judge + regression set) before deploy");
why.push("every stage can regress something the last stage fixed");
return { steps, why };
}
const cases: Array<[string, Signals]> = [
["support bot, prompting works", { promptingAndRagFail: false, haveDemonstrations: 0, havePreferencePairs: 0, verifiableReward: false, needsOwnBaseModel: false, gpuMemoryGB: 80, modelParamsB: 8 }],
["SQL generator with test DB", { promptingAndRagFail: true, haveDemonstrations: 20000, havePreferencePairs: 0, verifiableReward: true, needsOwnBaseModel: false, gpuMemoryGB: 80, modelParamsB: 8 }],
["brand-tone assistant, 70B", { promptingAndRagFail: true, haveDemonstrations: 5000, havePreferencePairs: 30000, verifiableReward: false, needsOwnBaseModel: false, gpuMemoryGB: 80, modelParamsB: 70 }],
["on-device summariser", { promptingAndRagFail: true, haveDemonstrations: 300, havePreferencePairs: 0, verifiableReward: false, needsOwnBaseModel: true, gpuMemoryGB: 24, modelParamsB: 3 }],
];
for (const [name, s] of cases) {
const plan = pick(s);
console.log(`\n# ${name}`);
plan.steps.forEach((st, i) => console.log(` ${i + 1}. ${st}\n why: ${plan.why[i]}`));
}Output:
# support bot, prompting works
1. prompting + RAG + evals
why: cheapest lever works; do not train yet
# SQL generator with test DB
1. SFT via LoRA
why: 20000 demos teach format/behaviour; full FT needs ~128 GB vs 80 GB
2. RL with verifiable rewards (GRPO-style, no reward model)
why: a program grader is harder to flatter than a learned reward model (still audit for test hacking)
3. eval gate (benchmarks + judge + regression set) before deploy
why: every stage can regress something the last stage fixed
# brand-tone assistant, 70B
1. SFT via QLoRA (4-bit base + LoRA)
why: 5000 demos teach format/behaviour; full FT needs ~1120 GB vs 80 GB
2. DPO-family preference tuning (offline)
why: pairs exist; no reward model or online sampling needed
3. eval gate (benchmarks + judge + regression set) before deploy
why: every stage can regress something the last stage fixed
# on-device summariser
1. collect or synthesise demos first (or distill from a stronger teacher)
why: SFT on a few hundred noisy demos mostly teaches noise
2. distill: teacher generates, smaller student is trained (SFT on outputs or logits)
why: cheaper serving at a known quality gap
3. eval gate (benchmarks + judge + regression set) before deploy
why: every stage can regress something the last stage fixedExpected# support bot, prompting works 1. prompting + RAG + evals why: cheapest lever works; do not train yet # SQL generator with test DB 1. SFT via LoRA why: 20000 demos teach format/behaviour; full FT needs ~128 GB vs 80 GB 2. RL with verifiable rewards (GRPO-style, no reward model) why: a program grader is harder to flatter than a learned reward model (still audit for test hacking) 3. eval gate (benchmarks + judge + regression set) before deploy why: every stage can regress something the last stage fixed # brand-tone assistant, 70B 1. SFT via QLoRA (4-bit base + LoRA) why: 5000 demos teach format/behaviour; full FT needs ~1120 GB vs 80 GB 2. DPO-family preference tuning (offline) why: pairs exist; no reward model or online sampling needed 3. eval gate (benchmarks + judge + regression set) before deploy why: every stage can regress something the last stage fixed # on-device summariser 1. collect or synthesise demos first (or distill from a stronger teacher) why: SFT on a few hundred noisy demos mostly teaches noise 2. distill: teacher generates, smaller student is trained (SFT on outputs or logits) why: cheaper serving at a known quality gap 3. eval gate (benchmarks + judge + regression set) before deploy why: every stage can regress something the last stage fixed
Press Run. Snippets must be self-contained — no network, files, or native modules.
Decision chart: which post-training technique?
Decisions
- 1
D1
- nextShip prompting, keep the eval suite, do not train
- nextStep 2: do you have 1k+ clean demonstrations?
- 2
Ship prompting, keep the eval suite, do not train
- ?
Step 2: do you have 1k+ clean demonstrations?
- nextCollect, synthesise or distill demonstrations first
- nextStep 3: does full fine-tuning fit your GPUs?
- 4
Collect, synthesise or distill demonstrations first
- ?
Step 3: does full fine-tuning fit your GPUs?
- nextSFT with full fine-tuning
- nextSFT with LoRA
- nextSFT with QLoRA
- 6
SFT with full fine-tuning
- nextStep 4: can a program grade the outputs?
- 7
SFT with LoRA
- nextStep 4: can a program grade the outputs?
- 8
SFT with QLoRA
- nextStep 4: can a program grade the outputs?
- ?
Step 4: can a program grade the outputs?
- nextGRPO-style RL with verifiable rewards
- nextStep 5: thousands of preference pairs?
- 10
GRPO-style RL with verifiable rewards
- nextStep 6: eval gate, then canary deploy
- ?
Step 5: thousands of preference pairs?
- nextDPO family offline
- nextReward model plus PPO
- nextKTO-style unpaired preference loss
- 12
DPO family offline
- nextStep 6: eval gate, then canary deploy
- 13
Reward model plus PPO
- nextStep 6: eval gate, then canary deploy
- 14
KTO-style unpaired preference loss
- nextStep 6: eval gate, then canary deploy
- 15
Step 6: eval gate, then canary deploy
What happens if you choose otherwise
- Skip SFT and go straight to preference tuning: the base model rarely produces well-formed answers, so pairs compare two bad completions and the reward signal is mostly noise.
- Use PPO when you only have a few thousand fixed pairs: you pay for online RL infrastructure but your reward model is the bottleneck; DPO on the same pairs is usually as good and far simpler.
- Use DPO for math or code with a checker available: you throw away the strongest signal you have; online RL against the checker keeps finding new failures to fix.
- Full fine-tune a 70B model for a tone change: a LoRA adapter would do it at a fraction of the memory, and you could serve many such adapters on one base.
- Ship without an eval gate: a gain on the target task hides a drop in safety refusals or multilingual quality that customers find first.
Pitfalls
- Treating the training loss as the success metric. SFT loss falls while the model memorises; DPO loss falls while the chosen answer becomes less likely. Track evals and chosen log-probs.
- Mismatched chat templates between training and serving. Different special tokens at inference silently degrade every answer.
- Reusing benchmark data in training sets, including paraphrases. Scores go up, capability does not.
- Over-optimising a reward model. Gains on the reward model keep rising after real quality peaks; this is measured in the reward over-optimisation literature.
- Forgetting that every stage changes the model's output distribution, which changes what the next stage's data should look like.
The cluster map
- Supervised Fine-Tuning (SFT) - Instruction Data, Chat Templates, Packing & Loss Masking: instruction data formats, chat templates and special tokens, assistant-only loss masking, padding vs packing, and the hyperparameters that cause forgetting.
- Reward Models & RLHF with PPO - Bradley-Terry, KL Penalty, Critics & Reward Hacking: Bradley-Terry reward models, PPO with a KL penalty, critics and GAE, the four-model cost, and reward hacking with its defenses.
- Preference Optimization - DPO, IPO, ORPO, SimPO, KTO & GRPO Compared with PPO: DPO from the RLHF objective, IPO, KTO, ORPO and SimPO, likelihood displacement, and GRPO's group-relative advantages compared with PPO.
- LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter Merging: LoRA's low-rank update, rank, alpha and target modules, QLoRA's NF4, double quantisation and paged optimizers, and merging vs keeping adapters.
- Distillation & LLM Evals - Teacher-Student Training, Benchmarks, LLM-as-Judge, Regression Gates & Contamination: sequence-level, logit-level and on-policy distillation, eval types, LLM-as-judge done right, regression gates and contamination checks.
Interview Q&A
Walk me through how a base model becomes a chat assistant.
Answer
Pretrain on next-token prediction; SFT on curated demonstrations formatted with a chat template and loss only on assistant tokens; collect preferences; optimise them with RLHF (reward model plus PPO with a KL penalty) or a DPO-family loss, or GRPO-style RL when outputs are verifiable; optionally distill into a smaller model; gate with benchmarks, judges, regression and safety slices; canary deploy.
Why not just do more SFT instead of RLHF or DPO?
Answer
SFT imitates a single target and cannot express "A is better than B." Many qualities are easier to compare than to write, and preference losses push away from specific bad behaviours. SFT also gives no direct way to penalise outputs you never wrote down.
PPO vs DPO in one minute?
Answer
Same objective (reward minus KL to a reference). PPO samples online, scores with a reward model, and uses a critic for advantages: powerful but four models and many knobs. DPO turns the objective into a classification-style loss on fixed pairs: no reward model, no sampling, cheap and stable, but limited to the pairs you have.
When would you pick GRPO?
Answer
When a program can grade answers, such as unit tests, exact numeric answers or a SQL result check. GRPO samples a group per prompt and uses the group mean as the baseline, so it needs no critic, and the verifiable reward is much harder to flatter than a learned reward model.
Your fine-tuned model scores higher on the target benchmark but users complain. What do you check?
Answer
Regression slices (safety, other languages, other task families), contamination of the benchmark, judge bias if an LLM judge was used, chat template mismatch in serving, and reward hacking signs such as longer answers or more hedging.
How do LoRA and QLoRA change the economics?
Answer
They freeze the base and train a low-rank update, so optimizer state is tiny; QLoRA also stores the base in 4-bit NF4. A 70B model that needs over 1 TB of training memory for full fine-tuning fits on one 80 GB GPU with QLoRA, before activations.
Why does SFT on a single gold answer fall short of preference losses?
Answer
It treats every other answer as equally wrong, while preference losses push specifically away from the bad behaviour you observed.
In the pipeline toy, why did DPO with beta=0.1 collapse to one answer?
Answer
Small beta needs a large log-ratio margin before the loss saturates, so the policy drifts further from the reference; beta=1.0 saturated early and kept some diversity.
Check yourself
Pick a model behaviour you want to change. Decide whether prompting plus retrieval is enough; if not, write which stage you would start with, the signal you would collect, and the regression slices your eval gate must protect.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: LLM Inference Runtime — Prefill, KV, Batching & Parallelism, Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism, Agent Reliability — Evals, Guardrails, Tracing & Failure Modes, RAG Failure Modes — Hallucination, Stale Indexes, Evals & Grounding, Structured Outputs / Constrained Decoding, LLM Inference in Production - Quantization, Prompt Caching, Multi-LoRA Serving & Capacity Planning.
Go Deeper
- Training language models to follow instructions with human feedback (InstructGPT)
- Direct Preference Optimization (DPO)
- DeepSeekMath: introduces GRPO
- LoRA: Low-Rank Adaptation of Large Language Models
- QLoRA: Efficient Finetuning of Quantized LLMs
- Tulu 3: open post-training recipes and data
- Hugging Face TRL documentation (SFT, DPO, GRPO, PPO trainers)