Reward Models & RLHF with PPO - Bradley-Terry, KL Penalty, Critics & Reward Hacking
Bradley-Terry reward models from pairwise preferences, PPO with KL penalty to the SFT reference, critic/GAE, four-models-in-memory cost, reward hacking and over-optimisation; runnable RM hacking demo + PPO step.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
How is a reward model trained?
Answer
On pairwise comparisons with the Bradley-Terry loss, usually from the SFT model with a scalar head.
L2
Why do only reward differences matter?
Answer
The loss depends on r(chosen) - r(rejected), so rewards have no absolute scale.
L3
Why the KL penalty?
Answer
The RM is accurate only near its training distribution; the KL keeps the policy close to the SFT reference.
L4
What does the critic do?
Answer
Predicts expected future reward per token so advantages measure better or worse than expected.
L5
Why clip the ratio?
Answer
It refuses extra credit for moves beyond about 20 percent, so one lucky sample cannot drag the policy far.
L6
What is reward hacking?
Answer
The policy exploits RM errors, for example padding answers because labelers preferred longer ones.
L7
Why did teams move to DPO or GRPO?
Answer
PPO needs four models, online generation and careful tuning.
Failure modes
No KL penalty
The policy collapses onto RM exploits while reward climbs.
KL coefficient too high
The policy barely moves and you get the SFT model back.
Stale reward model
It never saw the RL policy's new behaviours and scores them unreliably.
Misconceptions
A rising RM reward curve means quality is rising.
Plot a held-out eval next to it; Gao et al. showed gold reward rises then falls.
The reference model can be any checkpoint.
It must be exactly the SFT checkpoint you started from.
Human labels are clean ground truth.
InstructGPT reported about 73 percent inter-labeler agreement.
Interviewer traps
Rubrics that reward thoroughness.
They produce reward models that reward length; use length-controlled comparisons.
PPO without reward normalisation or value warm-up.
Normalise RM outputs and warm up the critic to avoid noisy advantages.
Design scenario
Same prompt for every reader.
Requirements
Recover real quality, keep the KL budget sane, and catch over-optimisation early next time.
Failure assumptions
- Labelers preferred longer answers.
- The RM was trained only on SFT outputs.
- No gold eval was tracked during training.
Constraints
- Fixed labeling budget per month.
- Four models must fit the training cluster.
Prompt
Your RLHF run's reward keeps climbing but users say answers got longer and more sycophantic. Diagnose and redesign.
API
Which comparisons, rubrics and length controls feed the reward model?
Data
Which metrics are logged per step: RM reward, KL, length, held-out evals?
Architecture
How are policy, reference, RM and critic placed, and when do you refresh the RM?
Overview
RLHF (reinforcement learning from human feedback) is the classic way to teach a model preferences that demonstrations cannot express. It has two learned components. A reward model (RM) is trained on pairwise comparisons ("answer A is better than B") to output a scalar score. Then the policy (the SFT model) is optimised with PPO to maximise that score minus a KL penalty that keeps it close to the SFT reference, because the reward model is only trustworthy near the data it was trained on. RLHF is powerful and online: the policy keeps generating fresh samples and learning from them. It is also the most expensive and fragile stage of post-training: four models in memory, many hyperparameters, and a constant risk of reward hacking.
The reward model
Data. Annotators (or an AI judge under a written rubric, as in RLAIF and Constitutional AI) see one prompt and two or more responses sampled from the current model and pick the better one. Ranking K responses gives up to K(K-1)/2 pairs per prompt.
Model. Usually the SFT model with its language-model head replaced by a scalar head that reads the final token's hidden state.
Loss (Bradley-Terry). Model the probability that the chosen answer wins as sigmoid(r(chosen) - r(rejected)) and minimise -log sigmoid(r_w - r_l). Only differences matter, so rewards have no absolute scale; teams often normalise RM outputs before RL.
| Reward source | Strength | Weakness | Typical use |
|---|---|---|---|
| Human pairwise preferences + RM | Captures subtle qualities | Slow, costly, noisy (InstructGPT reported about 73 percent inter-labeler agreement) | General assistants |
| AI feedback (RLAIF, constitution or rubric) | Cheap, scalable, consistent | Inherits judge biases (length, position, self-preference) | Harmlessness, style, scaling human labels |
| Verifiable / rule-based rewards | Precise, hard to fool | Only for checkable tasks | Math, code, formatting constraints |
| Process reward models (score each step) | Credit assignment for long reasoning | Expensive step-level labels | Multi-step reasoning research |
Bound the drift or let the policy chase the RM?
Prefer
KL penalty plus held-out evals
Keep the policy near the SFT reference and stop at the eval peak.
- At beta=0.5 the answer stayed 190 tokens with KL 0.3.
- True quality never moved in the toy, so the RM gain was the hack.
- Early stopping on a gold eval catches the peak.
Alternative
Optimise the RM with no leash
Treat the reward model as ground truth.
- At beta=0.0 answers grew to 2000 tokens and RM reward to 17.54.
- KL reached 684.5 while true quality stayed 0.25.
- Outputs become long, repetitive or bizarre.
One PPO-RLHF iteration
Diagram 1 condensed: sample, score, update, and the failure path.
- 1
Sample responses
The policy generates rollouts for a batch of prompts. - 2
Score
The reward model scores each response; a per-token KL to the reference is folded in. - 3
Estimate advantages
The critic and GAE say which tokens did better than expected. - 4
Clipped update
Ratios beyond about 20 percent get no extra credit. - 5
Reward hacking
RM reward climbs while held-out quality falls.
RLHF with PPO, step by step
Diagram 1: one PPO-RLHF iteration and the failure path.
Decisions
- 1
Step 1: sample a batch of prompts
- nextStep 2: policy generates responses (online rollouts)
- 2
Step 2: policy generates responses (online rollouts)
- nextStep 3: reward model scores each full response
- 3
Step 3: reward model scores each full response
- nextStep 4: reference model gives log-probs, per-token KL penalty is subtracted
- 4
Step 4: reference model gives log-probs, per-token KL penalty is subtracted
- nextStep 5: critic estimates values, GAE turns rewards into per-token advantages
- 5
Step 5: critic estimates values, GAE turns rewards into per-token advantages
- nextStep 6: PPO clipped update of policy, regression update of critic, a few epochs
- 6
Step 6: PPO clipped update of policy, regression update of critic, a few epochs
- nextStep 7: RM reward up AND held-out evals up AND KL in budget?
- ?
Step 7: RM reward up AND held-out evals up AND KL in budget?
- nextStep 1: sample a batch of prompts
- nextFailure path: RM reward keeps rising while real quality falls (reward hacking)
- 8
Failure path: RM reward keeps rising while real quality falls (reward hacking)
- nextStop at best checkpoint, raise KL coefficient, refresh RM with new comparisons on current policy outputs
- 9
Stop at best checkpoint, raise KL coefficient, refresh RM with new comparisons on current policy outputs
- nextStep 1: sample a batch of prompts
Lesson map
Reward Models & RLHF with PPO - Bradley-Terry, KL Penalty, Critics & Reward Hacking
Bradley-Terry reward models from pairwise preferences, PPO with KL penalty to the SFT reference, critic/GAE, four-models-in-memory cost, reward hacking and over-optimisation; runnable RM hacking demo + PPO step.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: sample a batch of prompts"] s2["Step 2: policy generates responses (online rollouts)"] s3["Step 3: reward model scores each full response"] s4["Step 4: reference model gives log-probs, per-token KL penalty is subtracted"] s5["Step 5: critic estimates values, GAE turns rewards into per-token advantages"] s6["Step 6: PPO clipped update of policy, regression update of critic, a few epochs"] s7["Step 7: RM reward up AND held-out evals up AND KL in budget?"] f1["Failure path: RM reward keeps rising while real quality falls (reward hacking)"] f2["Stop at best checkpoint, raise KL coefficient, refresh RM with new comparisons on current policy outputs"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s4 -->|continues| s5 s5 -->|continues| s6 s6 -->|continues| s7 s7 -->|continues| s1 s7 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s1
The objective PPO-RLHF maximises is E[ r_RM(x, y) ] - beta * KL( pi(y|x) || pi_ref(y|x) ). In practice the KL is folded into the reward per token: every token gets -beta * (log pi_old(token) - log pi_ref(token)), and the RM score is added on the last token.
Why a critic? Policy gradients need a baseline to reduce variance: "was this token better or worse than expected?" The critic (value model) predicts expected future reward at each token; GAE (generalised advantage estimation) blends one-step and multi-step estimates with a lambda parameter.
Why clipping? PPO's ratio pi_new / pi_old measures how much a token's probability changed since sampling. The clipped objective min(ratio * A, clip(ratio, 1-eps, 1+eps) * A) refuses extra credit for moves beyond about 20 percent, which keeps updates near the data that produced them. Without it, one lucky sample can drag the policy far off.
// One PPO-RLHF update computed by hand for a 5-token response.
// Pieces: per-token KL penalty folded into the reward, reward-model score on the last
// token, GAE advantages from a value (critic) model, and the clipped surrogate objective.
const beta = 0.1; // KL coefficient
const gamma = 1.0; // no discounting inside one response (common in RLHF)
const lam = 0.95; // GAE lambda
const eps = 0.2; // PPO clip range
const tokens = ["The", "answer", "is", "42", "<eos>"];
const logpOld = [-0.9, -1.2, -0.3, -2.0, -0.1]; // policy that sampled the response
const logpRef = [-0.8, -1.0, -0.3, -2.6, -0.1]; // frozen SFT reference
const logpNew = [-0.7, -1.1, -0.3, -1.2, -0.1]; // policy after a few gradient steps
const values = [0.50, 0.55, 0.60, 0.70, 0.90]; // critic's estimate V(s_t)
const rmScore = 1.4; // reward model score for the full response
// 1) Rewards: -beta * (log pi_old - log pi_ref) every token, plus the RM score at the end.
const rewards = tokens.map((_, t) => -beta * (logpOld[t] - logpRef[t]) + (t === tokens.length - 1 ? rmScore : 0));
// 2) GAE: delta_t = r_t + gamma V_{t+1} - V_t ; A_t = sum_l (gamma*lam)^l delta_{t+l}
const adv = new Array(tokens.length).fill(0);
let running = 0;
for (let t = tokens.length - 1; t >= 0; t--) {
const nextV = t + 1 < tokens.length ? values[t + 1] : 0; // terminal state value = 0
const delta = rewards[t] + gamma * nextV - values[t];
running = delta + gamma * lam * running;
adv[t] = running;
}
// 3) Clipped surrogate per token.
console.log("tok reward adv ratio unclipped clipped used");
let total = 0;
tokens.forEach((tok, t) => {
const ratio = Math.exp(logpNew[t] - logpOld[t]);
const unclipped = ratio * adv[t];
const clipped = Math.min(Math.max(ratio, 1 - eps), 1 + eps) * adv[t];
const used = Math.min(unclipped, clipped); // pessimistic bound
total += used;
console.log(`${tok.padEnd(8)}${rewards[t].toFixed(3).padStart(7)}${adv[t].toFixed(3).padStart(7)}${ratio.toFixed(2).padStart(7)}${unclipped.toFixed(3).padStart(10)}${clipped.toFixed(3).padStart(9)}${used.toFixed(3).padStart(7)}`);
});
console.log(`mean surrogate objective (maximise) = ${(total / tokens.length).toFixed(3)}`);
console.log("\nToken '42' moved from logp -2.0 to -1.2 (ratio 2.23 > 1.2): the clip caps its credit at 1.2x the");
console.log("advantage, so one lucky sample cannot drag the policy far in a single update.");
console.log("Four models in memory: policy, reference, reward model, critic. That cost is a big part of why DPO (no RM, no critic) and GRPO (no critic) exist.");Output:
tok reward adv ratio unclipped clipped used
The 0.010 0.744 1.22 0.909 0.893 0.893
answer 0.020 0.720 1.11 0.796 0.796 0.796
is 0.000 0.684 1.00 0.684 0.684 0.684
42 -0.060 0.615 2.23 1.369 0.738 0.738
<eos> 1.400 0.500 1.00 0.500 0.500 0.500
mean surrogate objective (maximise) = 0.722
Token '42' moved from logp -2.0 to -1.2 (ratio 2.23 > 1.2): the clip caps its credit at 1.2x the
advantage, so one lucky sample cannot drag the policy far in a single update.
Four models in memory: policy, reference, reward model, critic. That cost is a big part of why DPO (no RM, no critic) and GRPO (no critic) exist.Expectedtok reward adv ratio unclipped clipped used The 0.010 0.744 1.22 0.909 0.893 0.893 answer 0.020 0.720 1.11 0.796 0.796 0.796 is 0.000 0.684 1.00 0.684 0.684 0.684 42 -0.060 0.615 2.23 1.369 0.738 0.738 <eos> 1.400 0.500 1.00 0.500 0.500 0.500 mean surrogate objective (maximise) = 0.722 Token '42' moved from logp -2.0 to -1.2 (ratio 2.23 > 1.2): the clip caps its credit at 1.2x the advantage, so one lucky sample cannot drag the policy far in a single update. Four models in memory: policy, reference, reward model, critic. That cost is a big part of why DPO (no RM, no critic) and GRPO (no critic) exist.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The cost of PPO-RLHF
| Model in memory | Role | Trained? |
|---|---|---|
| Policy (actor) | Generates and is optimised | Yes |
| Reference | Anchors the KL penalty | No (frozen SFT) |
| Reward model | Scores responses | No (trained beforehand) |
| Critic (value model) | Baseline for advantages | Yes |
Plus generation in the loop: rollouts are autoregressive and dominate wall-clock time, which is why RLHF frameworks colocate a fast inference engine with the trainer. This is the overhead DPO removes entirely and GRPO halves (no critic).
Reward hacking and over-optimisation
The reward model is a proxy. Optimise any proxy hard enough and the policy finds directions where the proxy and the truth disagree (Goodhart's law). Gao et al. measured this: as KL from the reference grows, the RM score keeps increasing while the score from a stronger "gold" reward model rises and then falls.
Common hacks: longer answers (length bias in labels), sycophancy (agreeing with the user), confident hedging, excessive formatting, refusal patterns that the RM over-rewards.
"""Bradley-Terry reward model + reward hacking + the KL penalty, in pure stdlib.
Each response has hidden truth: correct or not. Humans can (noisily) tell, and prefer
correct, polite answers. The reward model (RM) only sees surface features it can
compute: [politeness, length_in_100_tokens]. In this synthetic data longer answers
tended to be the correct ones, so the RM learns "longer is better" as a shortcut.
A policy optimising that RM then pads answers: classic reward hacking.
"""
import math, random
random.seed(1)
def sigmoid(x): return 1 / (1 + math.exp(-x))
def sample_response():
correct = random.random() < 0.5
length = max(0.2, random.gauss(3.0 if correct else 1.5, 0.5)) # correlated with correctness
return {"correct": 1.0 if correct else 0.0, "polite": random.random(), "length": length}
def true_quality(x): # what humans actually value
return 2.0 * x["correct"] + 0.5 * x["polite"]
def feats(x): # what the RM can see
return [x["polite"], x["length"]]
# 1) Preference data: pairs labelled by noisy humans (Bradley-Terry on true quality).
pairs = []
for _ in range(3000):
a, b = sample_response(), sample_response()
pairs.append((a, b) if random.random() < sigmoid(true_quality(a) - true_quality(b)) else (b, a))
# 2) RM r(x) = w . feats(x), trained with the pairwise loss -log sigmoid(r(chosen) - r(rejected)).
w = [0.0, 0.0]
for _ in range(30):
for c, r in pairs:
d = [ci - ri for ci, ri in zip(feats(c), feats(r))]
g = 1 - sigmoid(sum(wi * di for wi, di in zip(w, d)))
w = [wi + 0.01 * g * di for wi, di in zip(w, d)]
acc = sum(sum(wi * (ci - ri) for wi, ci, ri in zip(w, feats(c), feats(r))) > 0 for c, r in pairs) / len(pairs)
print(f"RM weights [polite, length] = {[round(x, 2) for x in w]} held-in pair accuracy = {acc:.2f}")
# 3) The "policy" controls only the length of a WRONG answer (it cannot make it correct).
# Objective: RM reward - beta * KL. For Gaussians with equal sigma, KL = (mu - mu_ref)^2 / (2 sigma^2).
REF_LEN, SIGMA, MAX_LEN = 1.5, 0.5, 20.0
def kl(length): return (length - REF_LEN) ** 2 / (2 * SIGMA ** 2)
print(f"\n{'beta':<6}{'length':>10}{'RM reward':>11}{'KL':>8}{'true quality':>14}")
for beta in (0.0, 0.05, 0.5, 5.0):
grid = [x / 10 for x in range(2, int(MAX_LEN * 10) + 1)]
L = max(grid, key=lambda L: w[0] * 0.5 + w[1] * L - beta * kl(L))
x = {"correct": 0.0, "polite": 0.5, "length": L}
print(f"{beta:<6}{L*100:>8.0f}tk{w[0]*0.5 + w[1]*L:>11.2f}{kl(L):>8.1f}{true_quality(x):>14.2f}")
print("\nTrue quality never moves (the answer is still wrong) while RM reward climbs with length:")
print("that gap is reward hacking. The KL penalty bounds the drift; fixing the RM data")
print("(length-controlled pairs, ensembles, length penalties) removes the incentive.")Output:
RM weights [polite, length] = [0.39, 0.87] held-in pair accuracy = 0.69
beta length RM reward KL true quality
0.0 2000tk 17.54 684.5 0.25
0.05 580tk 5.22 37.0 0.25
0.5 190tk 1.84 0.3 0.25
5.0 150tk 1.49 0.0 0.25
True quality never moves (the answer is still wrong) while RM reward climbs with length:
that gap is reward hacking. The KL penalty bounds the drift; fixing the RM data
(length-controlled pairs, ensembles, length penalties) removes the incentive.| Defense | How it helps | Cost |
|---|---|---|
| KL penalty to the reference | Bounds how far the policy can drift into RM blind spots | Limits achievable gains |
| Early stopping on held-out evals / gold RM | Catches the peak before over-optimisation | Needs a trusted eval |
| Length-controlled comparisons, length penalty | Removes the easiest shortcut | Data work |
| RM ensembles, uncertainty penalties | Hacks rarely fool all members | More compute |
| Iterated RLHF: new comparisons on current policy outputs | RM learns the policy's new failure modes | Ongoing labeling |
Decision chart: is PPO-RLHF the right tool?
Decisions
- 1
D1
- nextUse verifiable rewards, consider GRPO (no RM, no critic)
- nextStep 2: do you need online improvement beyond a fixed pair set?
- 2
Use verifiable rewards, consider GRPO (no RM, no critic)
- ?
Step 2: do you need online improvement beyond a fixed pair set?
- nextDPO family on the pairs, far simpler
- nextStep 3: can you host policy, reference, RM and critic plus a rollout engine?
- 4
DPO family on the pairs, far simpler
- ?
Step 3: can you host policy, reference, RM and critic plus a rollout engine?
- nextIterative DPO: sample, label with RM or judge, DPO, repeat
- nextStep 4: is the reward model validated on held-out pairs and stress tests?
- 6
Iterative DPO: sample, label with RM or judge, DPO, repeat
- ?
Step 4: is the reward model validated on held-out pairs and stress tests?
- nextFix the RM first, RL amplifies RM errors
- nextPPO-RLHF with KL budget, early stopping and gold evals
- 8
Fix the RM first, RL amplifies RM errors
- 9
PPO-RLHF with KL budget, early stopping and gold evals
What happens if you choose otherwise
- No KL penalty: the policy collapses onto RM exploits within a few hundred steps; outputs become long, repetitive or bizarre while reward climbs.
- KL coefficient too high: the policy barely moves; you paid for RLHF and got the SFT model.
- Stale reward model: trained on SFT outputs, it has never seen the RL policy's new behaviours and scores them unreliably.
- PPO without reward normalisation or value warm-up: noisy advantages and unstable training.
Pitfalls
- Reading RM reward curves as quality curves. Always plot a held-out eval next to them.
- Annotator guidelines that reward thoroughness produce RMs that reward length.
- Pairs from a single model family make the RM blind to other styles.
- Forgetting that the reference model must be exactly the SFT checkpoint you started from.
Interview Q&A
How is a reward model trained?
Answer
On pairwise comparisons with a Bradley-Terry loss, -log sigmoid(r_chosen - r_rejected), usually initialised from the SFT model with a scalar head. Only reward differences are meaningful.
Why the KL penalty in RLHF?
Answer
The RM is accurate only near its training distribution. The KL term keeps the policy close to the SFT reference, limiting reward hacking and preserving fluency and diversity.
What does the critic do in PPO?
Answer
It predicts expected future reward per token so advantages measure "better or worse than expected," cutting gradient variance. GAE blends multi-step returns for a bias-variance trade-off.
What is reward hacking? Give an example and a fix.
Answer
The policy exploits RM errors to get high reward without real quality, for example padding answers because labelers preferred longer ones. Fixes: KL penalty, early stopping on gold evals, length-controlled data, RM ensembles, refreshing the RM on current policy outputs.
Why did many teams move from PPO to DPO or GRPO?
Answer
PPO needs four models, online generation and careful tuning. DPO removes the RM, critic and sampling for offline pairs; GRPO keeps online RL but removes the critic, and pairs naturally with verifiable rewards.
What are RLAIF and Constitutional AI in reward modeling?
Answer
An AI judge under a written rubric or constitution replaces human annotators: cheap and consistent, but it inherits judge biases.
Why do RLHF frameworks colocate a fast inference engine with the trainer?
Answer
Rollouts are autoregressive and dominate wall-clock time.
How many pairs does ranking K responses give?
Answer
Up to K(K-1)/2 pairs per prompt.
Check yourself
Take a reward or quality metric you optimise and list the cheapest ways a system could raise it without real improvement. For each, name the leash or eval that would catch it.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism, Algorithm Selection Playbook — Metrics, Baselines & Failure Modes.
Go Deeper
- Proximal Policy Optimization Algorithms (PPO)
- High-Dimensional Continuous Control Using Generalized Advantage Estimation (GAE)
- Training language models to follow instructions with human feedback (InstructGPT)
- Scaling Laws for Reward Model Overoptimization
- Constitutional AI: Harmlessness from AI Feedback
- Hugging Face TRL documentation