Preference Optimization - DPO, IPO, ORPO, SimPO, KTO & GRPO Compared with PPO
DPO derivation from the RLHF objective; IPO, KTO, ORPO, SimPO variants (reference-free, unpaired data); GRPO group-relative advantages with verifiable rewards; all compared with PPO by signal, cost, failure mode; runnable losses + GRPO advantages.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is DPO's key insight?
Answer
The optimal KL-regularised policy lets the reward be written as beta times log(pi / pi_ref) up to a constant.
L2
What does beta control?
Answer
The strength of the implicit KL anchor: small beta drifts further before the loss saturates.
L3
What is likelihood displacement?
Answer
The margin grows while both chosen and rejected log-probs fall.
L4
When pick KTO?
Answer
When you only have unpaired good or bad labels such as thumbs from logs.
L5
What makes ORPO and SimPO reference-free?
Answer
ORPO adds an odds-ratio term to the SFT loss; SimPO uses a length-normalised margin.
L6
How does GRPO avoid a critic?
Answer
It uses the group's mean and std reward as the baseline.
L7
When does a GRPO group give no gradient?
Answer
When every sample is right or every sample is wrong.
Failure modes
Likelihood displacement
DPO loss falls while the chosen answer becomes less likely.
GRPO with a weak grader
The policy satisfies the grader without solving, for example special-casing test inputs.
Tiny beta on clean pairs
The policy chases ever-larger margins and collapses diversity.
Misconceptions
Falling DPO loss means better outputs.
Check chosen log-probs; the toy's displacement row had lower loss and a less likely chosen answer.
DPO and PPO optimise different objectives.
Both optimise reward minus a KL leash; DPO rewrites it in closed form.
ORPO needs an SFT stage first.
It merges SFT and preference tuning, at the cost of no reference safety net.
Interviewer traps
Comparing methods at different beta scales or data.
Hold beta and data comparable before attributing differences to the method.
Using DPO for math or code with a checker available.
Online RL against the checker uses the strongest signal you have.
Design scenario
Same prompt for every reader.
Requirements
Better preferences on support without drift, better accuracy on math, and monitoring that catches displacement or grader gaming.
Failure assumptions
- Support pairs were sampled from an older model.
- Chosen answers are often longer.
- Many math prompts are all-correct or all-wrong.
Constraints
- Limited GPU memory for extra models.
- A weekly retraining slot.
Prompt
You have 30,000 offline preference pairs for a support assistant and a separate math tutor with exact-answer checks. Choose methods for each.
API
What data shape (pairs, unpaired labels, prompts plus grader) does each method consume?
Data
Which metrics catch displacement and zero-signal groups?
Architecture
Which method, reference model and sampling loop runs for each product?
Overview
DPO (Direct Preference Optimization) showed that the RLHF objective, "maximise reward minus a KL leash to a reference model," has a closed-form optimal policy, and that the reward can be rewritten in terms of the policy itself: r(x, y) = beta * log( pi(y|x) / pi_ref(y|x) ) + const. Plug that into the Bradley-Terry preference loss and the reward model disappears: you train the policy directly on (prompt, chosen, rejected) triples with a supervised-looking loss. A family of variants (IPO, KTO, ORPO, SimPO) changes the loss shape, drops the reference model, or works with unpaired thumbs up/down data. GRPO goes the other way: it keeps online RL but drops PPO's critic, and pairs naturally with verifiable rewards. This page compares all of them against PPO by signal, cost and failure mode.
The DPO loss, read slowly
For each pair, compute summed token log-probs of the chosen (w) and rejected (l) responses under the policy and under the frozen reference:
loss = -log sigmoid( beta * [ (log pi(w) - log pi_ref(w)) - (log pi(l) - log pi_ref(l)) ] )
- The bracket is the implicit reward margin: how much more the policy has raised the chosen answer than the rejected one, relative to the reference.
- beta plays the role of the KL coefficient. Small beta: the loss keeps pushing until the log-ratio margin is large, so the policy drifts further. Large beta: the loss saturates early, the policy stays close.
- The gradient on each pair is weighted by
sigmoid(-beta * margin): pairs the model already ranks correctly contribute little, wrong pairs contribute a lot.
The family, compared
| Method | Data | Reference model? | Loss idea | Why pick it | Known weakness |
|---|---|---|---|---|---|
| PPO-RLHF | Prompts + RM (from pairs), online samples | Yes (KL) | Maximise RM reward minus KL with a critic | Online, can keep improving | 4 models, unstable, expensive |
| DPO | Pairs, offline | Yes | Logistic loss on implicit reward margin | Simple, stable, cheap | Off-policy staleness, likelihood displacement |
| IPO | Pairs | Yes | Squared loss toward a fixed target margin | Does not overfit deterministic preferences | Target margin is another hyperparameter |
| KTO | Unpaired good / bad labels | Yes | Prospect-theory value per example | Uses thumbs up/down logs, no pairing needed | Weaker signal per example |
| ORPO | Pairs | No | SFT loss + odds-ratio preference term in one stage | Merges SFT and preference tuning, less memory | Less control over drift (no reference) |
| SimPO | Pairs | No | Length-normalised log-prob margin with a target gap | No reference forward pass, reduces length exploitation | Sensitive to beta and gamma |
| GRPO | Prompts + a reward function, online groups | Optional KL to reference | Group-relative advantage, PPO-style clipping, no critic | Verifiable rewards (math, code), reasoning | Sampling cost, zero-variance groups |
Offline DPO or online GRPO?
Prefer
GRPO when answers are checkable
Sample G completions per prompt and use group-relative advantages.
- Rare successes got large advantages: 2.65 on the hard 7^5 mod 13 prompt.
- No value model to train or store.
- Only 2 of 4 prompts gave signal; all-same groups give zero advantage.
Alternative
DPO on fixed pairs
A logistic loss on the implicit reward margin.
- Two models, and reference log-probs can be precomputed.
- At margin 12 with beta=0.1 the loss was 0.263 with gradient weight 0.231.
- Pairs go stale as the policy moves; displacement lowered loss to 0.105 while the chosen answer got less likely.
One GRPO step
Diagram 1 condensed: group sampling with a verifiable reward, and the failure path.
- 1
Sample a group
G completions per prompt from the current policy. - 2
Grade
A program checks each answer. - 3
Group-relative advantage
Reward minus group mean, over group std. - 4
Clipped update
PPO-style clipping, usually with a KL to the reference. - 5
Zero-variance or gamed groups
All-same groups give no gradient; a weak grader gets gamed.
"""DPO-family losses on sequence log-probs, plus the failure mode they share.
Inputs per preference pair: summed token log-probs of the chosen (w) and rejected (l)
responses under the policy and under the frozen reference, plus token counts.
DPO : -log sigmoid(beta * [(lp_w - ref_w) - (lp_l - ref_l)])
IPO : ([(lp_w - ref_w) - (lp_l - ref_l)] - 1/(2*tau))^2 (squared loss toward a target margin; overshooting is penalised)
SimPO : -log sigmoid(beta * (lp_w/n_w - lp_l/n_l) - gamma) (no reference model, length-normalised)
ORPO : NLL(chosen) - lambda * log sigmoid(log odds_w - log odds_l), odds = p/(1-p), p = exp(lp/n)
"""
import math
def logsig(x): return -math.log1p(math.exp(-x)) if x > -30 else x
def sig(x): return 1 / (1 + math.exp(-x))
def dpo(p, beta=0.1):
m = (p["lp_w"] - p["ref_w"]) - (p["lp_l"] - p["ref_l"])
return -logsig(beta * m), m, sig(-beta * m) # loss, margin, gradient weight
def ipo(p, tau=0.1):
m = (p["lp_w"] - p["ref_w"]) - (p["lp_l"] - p["ref_l"])
return (m - 1 / (2 * tau)) ** 2
def simpo(p, beta=2.0, gamma=0.5):
return -logsig(beta * (p["lp_w"] / p["n_w"] - p["lp_l"] / p["n_l"]) - gamma)
def orpo(p, lam=0.1):
def log_odds(lp, n):
avg = lp / n
return avg - math.log1p(-math.exp(avg))
nll = -p["lp_w"] / p["n_w"]
return nll - lam * logsig(log_odds(p["lp_w"], p["n_w"]) - log_odds(p["lp_l"], p["n_l"]))
before = dict(lp_w=-40.0, lp_l=-42.0, ref_w=-40.0, ref_l=-42.0, n_w=20, n_l=20) # policy == reference
after_good = dict(before, lp_w=-36.0, lp_l=-50.0) # chosen up, rejected down
after_displaced = dict(before, lp_w=-46.0, lp_l=-70.0) # BOTH down, rejected more: margin still grows
print(f"{'state':<18}{'margin':>7}{'DPO b=0.1':>10}{'grad w':>8}{'IPO':>8}{'SimPO':>8}{'ORPO':>8}")
for name, p in [("start (pi=ref)", before), ("healthy update", after_good), ("displacement", after_displaced)]:
l, m, g = dpo(p)
print(f"{name:<18}{m:>7.1f}{l:>10.3f}{g:>8.3f}{ipo(p):>8.2f}{simpo(p):>8.3f}{orpo(p):>8.3f}")
print("\nbeta sweep at the healthy margin of 12 (how hard DPO still pushes):")
for beta in (0.01, 0.1, 0.5):
l, m, g = dpo(after_good, beta)
print(f" beta={beta:<5} loss={l:.3f} gradient weight sigmoid(-beta*m)={g:.3f}")
print("\nFailure path ('displacement'): DPO loss went DOWN although the chosen answer became LESS")
print("likely (-40 -> -46). Optimising a relative margin can drain probability from both answers to")
print("unseen text. Watch chosen log-prob, not just loss; mitigations: SFT term on chosen (as in ORPO/RPO),")
print("larger beta, fresher on-policy pairs. Note IPO and ORPO both score 'displacement' WORSE than the")
print("healthy update: IPO penalises overshooting its target margin 1/(2*tau)=5, ORPO's NLL term rises.")Output:
state margin DPO b=0.1 grad w IPO SimPO ORPO
start (pi=ref) 0.0 0.693 0.500 25.00 0.854 2.064
healthy update 12.0 0.263 0.231 49.00 0.341 1.837
displacement 22.0 0.105 0.100 289.00 0.139 2.325
beta sweep at the healthy margin of 12 (how hard DPO still pushes):
beta=0.01 loss=0.635 gradient weight sigmoid(-beta*m)=0.470
beta=0.1 loss=0.263 gradient weight sigmoid(-beta*m)=0.231
beta=0.5 loss=0.002 gradient weight sigmoid(-beta*m)=0.002
Failure path ('displacement'): DPO loss went DOWN although the chosen answer became LESS
likely (-40 -> -46). Optimising a relative margin can drain probability from both answers to
unseen text. Watch chosen log-prob, not just loss; mitigations: SFT term on chosen (as in ORPO/RPO),
larger beta, fresher on-policy pairs. Note IPO and ORPO both score 'displacement' WORSE than the
healthy update: IPO penalises overshooting its target margin 1/(2*tau)=5, ORPO's NLL term rises.Why DPO usually beats PPO for teams, and when it does not
DPO wins on engineering: two models (policy + frozen reference, and the reference log-probs can be precomputed), no rollouts, no reward model, ordinary supervised training infrastructure, far fewer knobs.
PPO (or online methods) win on signal: DPO learns only from the pairs you collected, which were sampled from an older model. As the policy changes, those pairs say less and less about its current mistakes. Online methods sample from the current policy every step. Practical middle ground: iterative / online DPO, where you periodically sample from the current policy, label new pairs with a reward model or judge, and run DPO again.
What happens if you pick DPO for a checkable task: you leave the strongest signal (the checker) unused and learn from a static snapshot.
Likelihood displacement
DPO optimises a relative quantity. The margin can grow while both chosen and rejected log-probs fall, with probability mass leaking to text that appears in neither. This has been documented as likelihood displacement and can even reduce the probability of good responses. The toy output above shows it: loss down, chosen log-prob down. Mitigations: monitor chosen log-probs, add an SFT term on the chosen response (as ORPO and RPO-style variants do), raise beta, use fresher on-policy pairs, and avoid pairs where chosen and rejected are nearly identical.
GRPO: online RL without a critic
GRPO (introduced in DeepSeekMath and used at scale for DeepSeek-R1-style reasoning training) samples a group of G completions per prompt, scores each with a reward function, and uses (r_i - mean(r)) / std(r) inside the group as every token's advantage. The group mean replaces the critic's baseline. The policy update reuses PPO's clipped ratio and typically adds a KL term to the reference.
// GRPO-style group-relative advantages with a verifiable reward.
// For each prompt, sample G completions, score each with a program (here: does the
// final answer equal the expected one?), and use (r - mean) / std inside the group as
// the advantage. No critic network: the group mean IS the baseline.
function lcg(seed: number): () => number {
let s = seed >>> 0;
return () => ((s = (Math.imul(s, 1103515245) + 12345) >>> 0) / 2 ** 32);
}
const rnd = lcg(42);
type Prompt = { q: string; answer: number; pSolve: number }; // pSolve = current policy's success rate
const prompts: Prompt[] = [
{ q: "17*3", answer: 51, pSolve: 0.95 }, // too easy: almost every sample is right
{ q: "sum of primes < 30", answer: 129, pSolve: 0.5 },
{ q: "7^5 mod 13", answer: 11, pSolve: 0.1 }, // hard: rare success
{ q: "unsolvable for now", answer: -1, pSolve: 0.0 },
];
const G = 8;
function advantages(rewards: number[]): number[] {
const mean = rewards.reduce((a, b) => a + b, 0) / rewards.length;
const std = Math.sqrt(rewards.reduce((a, r) => a + (r - mean) ** 2, 0) / rewards.length);
return rewards.map((r) => (std < 1e-8 ? 0 : (r - mean) / std)); // zero-variance group -> no signal
}
let useful = 0;
for (const p of prompts) {
// Verifiable reward: 1 if the sampled final answer is correct, else 0 (a real grader parses the output).
const rewards = Array.from({ length: G }, () => (rnd() < p.pSolve ? 1 : 0));
const adv = advantages(rewards);
const informative = adv.some((a) => a !== 0);
if (informative) useful++;
console.log(`${p.q.padEnd(20)} rewards=[${rewards.join("")}] adv=[${adv.map((a) => a.toFixed(2)).join(", ")}]${informative ? "" : " <- no gradient"}`);
}
console.log(`\n${useful}/${prompts.length} prompts produced a learning signal this step.`);
console.log("Rare successes on hard prompts get large positive advantages (they are surprising relative to the group).");
console.log("All-correct or all-wrong groups give zero advantage: filter or re-balance prompts by difficulty (curriculum).");
console.log("Compared with PPO: no value model to train or store; cost moves to sampling G completions per prompt.");Output:
17*3 rewards=[11111111] adv=[0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00] <- no gradient
sum of primes < 30 rewards=[11010011] adv=[0.77, 0.77, -1.29, 0.77, -1.29, -1.29, 0.77, 0.77]
7^5 mod 13 rewards=[00100000] adv=[-0.38, -0.38, 2.65, -0.38, -0.38, -0.38, -0.38, -0.38]
unsolvable for now rewards=[00000000] adv=[0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00] <- no gradient
2/4 prompts produced a learning signal this step.
Rare successes on hard prompts get large positive advantages (they are surprising relative to the group).
All-correct or all-wrong groups give zero advantage: filter or re-balance prompts by difficulty (curriculum).
Compared with PPO: no value model to train or store; cost moves to sampling G completions per prompt.Expected17*3 rewards=[11111111] adv=[0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00] <- no gradient sum of primes < 30 rewards=[11010011] adv=[0.77, 0.77, -1.29, 0.77, -1.29, -1.29, 0.77, 0.77] 7^5 mod 13 rewards=[00100000] adv=[-0.38, -0.38, 2.65, -0.38, -0.38, -0.38, -0.38, -0.38] unsolvable for now rewards=[00000000] adv=[0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00, 0.00] <- no gradient 2/4 prompts produced a learning signal this step. Rare successes on hard prompts get large positive advantages (they are surprising relative to the group). All-correct or all-wrong groups give zero advantage: filter or re-balance prompts by difficulty (curriculum). Compared with PPO: no value model to train or store; cost moves to sampling G completions per prompt.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Diagram 1: one GRPO step with a verifiable reward, and the failure path.
Decisions
- 1
Step 1: pick a batch of prompts with known answers or tests
- nextStep 2: sample G completions per prompt from the current policy
- 2
Step 2: sample G completions per prompt from the current policy
- nextStep 3: grade each completion with a program (exact answer, unit tests, format check)
- 3
Step 3: grade each completion with a program (exact answer, unit tests, format check)
- nextStep 4: advantage = (reward - group mean) / group std, shared by every token of that completion
- 4
Step 4: advantage = (reward - group mean) / group std, shared by every token of that completion
- nextStep 5: PPO-style clipped update plus KL to the reference
- 5
Step 5: PPO-style clipped update plus KL to the reference
- nextStep 6: accuracy on held-out problems rising, response length sane?
- ?
Step 6: accuracy on held-out problems rising, response length sane?
- nextStep 1: pick a batch of prompts with known answers or tests
- nextFailure path: groups all-correct or all-wrong (no signal), or the policy games the checker
- 7
Failure path: groups all-correct or all-wrong (no signal), or the policy games the checker
- nextRebalance prompt difficulty, harden the grader, add format and length checks
- 8
Rebalance prompt difficulty, harden the grader, add format and length checks
- nextStep 1: pick a batch of prompts with known answers or tests
Lesson map
Preference Optimization - DPO, IPO, ORPO, SimPO, KTO & GRPO Compared with PPO
DPO derivation from the RLHF objective; IPO, KTO, ORPO, SimPO variants (reference-free, unpaired data); GRPO group-relative advantages with verifiable rewards; all compared with PPO by signal, cost, failure mode; runnable losses + GRPO advantages.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: pick a batch of prompts with known answers or tests"] s2["Step 2: sample G completions per prompt from the current policy"] s3["Step 3: grade each completion with a program (exact answer, unit tests, format check)"] s4["Step 4: advantage = (reward - group mean) / group std, shared by every token of that completion"] s5["Step 5: PPO-style clipped update plus KL to the reference"] s6["Step 6: accuracy on held-out problems rising, response length sane?"] f1["Failure path: groups all-correct or all-wrong (no signal), or the policy games the checker"] f2["Rebalance prompt difficulty, harden the grader, add format and length checks"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s4 -->|continues| s5 s5 -->|continues| s6 s6 -->|continues| s1 s6 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s1
Decision chart: which preference method?
Decisions
- 1
D1
- nextGRPO-style online RL with verifiable rewards
- nextStep 2: what labels do you have?
- 2
GRPO-style online RL with verifiable rewards
- ?
Step 2: what labels do you have?
- nextKTO
- nextStep 3: memory too tight for a reference model?
- 4
KTO
- ?
Step 3: memory too tight for a reference model?
- nextORPO (merge with SFT) or SimPO
- nextStep 4: preferences near-deterministic or tiny dataset?
- 6
ORPO (merge with SFT) or SimPO
- ?
Step 4: preferences near-deterministic or tiny dataset?
- nextIPO, or DPO with larger beta
- nextDPO, then iterate with fresh on-policy pairs
- 8
IPO, or DPO with larger beta
- 9
DPO, then iterate with fresh on-policy pairs
- nextStep 5: need continual online gains and can afford 4 models?
- ?
Step 5: need continual online gains and can afford 4 models?
- nextPPO-RLHF with a validated reward model
- nextStay with iterative DPO
- 11
PPO-RLHF with a validated reward model
- 12
Stay with iterative DPO
What happens if you choose otherwise
- DPO with tiny beta on clean pairs: the policy chases ever-larger margins, collapses diversity and may lower chosen likelihoods.
- ORPO on a model that has not been instruction-tuned at all: works by design (it merges SFT), but you lose the safety net of a reference; monitor drift on general evals.
- GRPO with a weak grader: the policy learns to satisfy the grader (print the expected answer format without solving, special-case test inputs).
- PPO for a dataset of 10k static pairs: heavy infrastructure for a signal DPO can extract as well.
Pitfalls
- Pairs where chosen and rejected differ by a few tokens produce large gradients on those tokens and encourage displacement elsewhere.
- Length bias: chosen answers are often longer, so DPO learns "longer is better." SimPO's length normalisation and length-balanced data help.
- Reusing the same pairs for many epochs overfits the implicit reward.
- Comparing methods with different beta scales or data as if results were method differences.
Interview Q&A
Derive the intuition behind DPO in two sentences.
Answer
The KL-regularised RLHF objective has an optimal policy proportional to pi_ref * exp(r / beta), so the reward equals beta * log(pi / pi_ref) up to a constant. Substituting into the Bradley-Terry loss gives a loss on policy log-probs alone, with no reward model or sampling.
What does beta control in DPO?
Answer
The strength of the implicit KL anchor. The loss saturates once beta * margin is large, so small beta lets the policy drift further from the reference before the gradient vanishes; large beta keeps it close.
DPO loss is falling but outputs got worse. What could be happening?
Answer
Likelihood displacement: both chosen and rejected log-probs fall while the margin grows. Check chosen log-probs, add an SFT term, raise beta, refresh pairs on-policy.
How does GRPO avoid needing a critic?
Answer
It samples several completions per prompt and uses the group's mean (and std) reward as the baseline, so the advantage is "better or worse than siblings for the same prompt."
When would you choose KTO or ORPO?
Answer
KTO when you only have unpaired good/bad feedback such as thumbs from production logs. ORPO when you want one stage combining SFT and preference learning without holding a reference model in memory.
What is iterative or online DPO?
Answer
Periodically sample from the current policy, label new pairs with a reward model or judge, and run DPO again.
How does the DPO gradient weight each pair?
Answer
By sigmoid(-beta * margin): pairs already ranked correctly contribute little, wrong pairs contribute a lot.
Why do pairs that differ by a few tokens cause trouble?
Answer
They put large gradients on those tokens and encourage displacement elsewhere.
Check yourself
List the preference data you actually have (pairs, thumbs, or a grader). Pick a method from the family table, name its known weakness, and write the metric you would log to catch it.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Agent Reliability — Evals, Guardrails, Tracing & Failure Modes, Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism.
Go Deeper
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- A General Theoretical Paradigm to Understand Learning from Human Preferences (IPO)
- KTO: Model Alignment as Prospect Theoretic Optimization
- ORPO: Monolithic Preference Optimization without Reference Model
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- DeepSeekMath (GRPO)
- DeepSeek-R1: incentivizing reasoning via RL
- Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization