Distillation & LLM Evals - Teacher-Student Training, Benchmarks, LLM-as-Judge, Regression Gates & Contamination
Sequence-level vs logit-level vs on-policy distillation and reasoning-trace distillation; benchmarks, LLM-as-judge biases, regression gates with per-slice thresholds and CIs, contamination checks; runnable distill/contamination + regression gate.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Sequence-level vs logit-level distillation?
Answer
Teacher text through any API vs teacher distributions with white-box access and the same vocabulary.
L2
Why use temperature?
Answer
A softened teacher distribution exposes the similarity structure among wrong answers.
L3
Forward vs reverse KL?
Answer
Forward spreads to cover every teacher mode; reverse commits to modes the teacher likes.
L4
How do you make an LLM judge trustworthy?
Answer
Rubric, pairwise comparisons, position swaps, length control, a different judge model, measured human agreement.
L5
What belongs in a regression gate?
Answer
A paired CI on overall quality, no-regression tolerances per slice, hard floors, and judge sanity checks.
L6
Why are averages dangerous?
Answer
A better average hides a protected-slice regression.
L7
How do you detect contamination?
Answer
n-gram overlap, embedding similarity, behavioural probes, and fresh or private held-out sets.
Failure modes
Shipping a student without a gap measurement
Users find the slices where the student is worse.
Single-number eval
A safety regression hides behind a better average.
Paraphrased contamination
Items pass the n-gram check but are the same problem.
Misconceptions
A strong LLM judge is unbiased.
Judges show position, length and self-preference bias.
Public benchmarks are enough.
They saturate and contaminate; a private regression set catches real breakage.
Small deltas on a few hundred items are real.
Without confidence intervals they may be noise.
Interviewer traps
Using the same model family as teacher, student and judge.
Use a different judge and calibrate it against human labels.
Evaluating with different decoding settings than production.
Match temperature, max tokens and chat template.
Design scenario
Same prompt for every reader.
Requirements
A known teacher-student gap per slice, no safety regression, and a gate that blocks bad releases automatically.
Failure assumptions
- The teacher is only available through an API.
- An LLM judge flips verdicts when answer order is swapped.
- A benchmark leaked into synthetic training data.
Constraints
- Licence terms must allow distillation.
- Releases happen weekly.
Prompt
Distill a 70B assistant into an 8B student for a cost-sensitive product and design its release gate.
API
Which teacher outputs (text, top-k logits) and judge prompts are available?
Data
Which slices, versioned eval sets and contamination checks back the gate?
Architecture
Where does the gate sit between candidate checkpoint, canary and production?
Overview
Two stages close the post-training pipeline. Distillation moves capability from a large, expensive teacher into a smaller, cheaper student, either by training the student on the teacher's generated text (sequence-level, black-box) or on the teacher's full probability distributions (logit-level, white-box), offline or on the student's own samples. Evals decide whether any checkpoint, distilled or not, may ship: offline benchmarks for capability, LLM-as-judge for open-ended quality, regression gates with per-slice thresholds and confidence intervals, and contamination checks so the numbers mean something. This page treats distillation as a cost lever with a known quality gap and evals as the release process that keeps every other stage honest. For evals of agents (trajectories, tool calls, guardrails) and RAG (grounding, citations), see the existing pages linked under Related; this page covers model-level evals.
Distillation, compared
| Approach | Needs from the teacher | Student learns from | Strength | Weakness |
|---|---|---|---|---|
| Sequence-level (generate then SFT) | Text outputs only (works through an API) | Teacher's sampled answers | Simple, black-box, scales | Loses the teacher's uncertainty; licence terms may forbid it |
| Logit-level (soft targets, KL) | Full or top-k token distributions, same tokenizer | Teacher probabilities, often at temperature T | Much more signal per token ("dark knowledge") | Needs white-box access and matching vocabularies |
| On-policy distillation (MiniLLM, GKD) | Teacher log-probs on student samples | Teacher's judgement of what the student actually generates | Fixes exposure bias, student learns to recover from its own mistakes | Online sampling cost |
| Reasoning-trace distillation | Long chain-of-thought outputs | Teacher's step-by-step solutions | Transfers reasoning style to small models efficiently | Student imitates surface form, long outputs cost tokens |
Gate on the overall average or on every protected slice?
Prefer
Paired CI plus per-slice no-regression rules
Rules fixed before the results are seen.
- Overall delta was +4.3pp with a CI excluding zero.
- Safety refusals regressed -6.5pp, so the gate said FAIL.
- Judge flips on 30% of swapped pairs were counted as ties.
Alternative
One overall number
Ship when the average improves.
- The +4.3pp average would have shipped.
- The safety regression reaches users first.
- Looks like a win until incidents arrive.
From candidate checkpoint to production
Diagram 1 condensed: the release pipeline and its failure path.
- 1
Contamination check
n-gram overlap plus embedding similarity against training data. - 2
Offline evals
Benchmarks, execution checks and calibrated pairwise judges. - 3
Regression gate
Paired CI overall, no protected slice may regress, hard floors. - 4
Canary
Real user outcomes on a small share of traffic. - 5
Gate fails
Block the release and fix the regressing slice.
Forward vs reverse KL. Classic logit distillation minimises KL(teacher || student) (forward): the student is penalised wherever the teacher has mass and the student does not, so it spreads to cover every teacher mode, which a small student may do badly, producing bland averages. Reverse KL KL(student || teacher), used by MiniLLM, penalises the student for putting mass where the teacher has none, so it commits to modes the teacher likes. GKD generalises this with divergence choices on student-generated samples.
Temperature. Dividing logits by T > 1 softens the teacher's distribution so the ranking among wrong answers (cat is closer to lynx than to truck) becomes visible to the loss; the loss is multiplied by T squared to keep gradient scale comparable (Hinton et al.).
"""Two eval-adjacent tools in pure stdlib:
(1) knowledge distillation losses: temperature-softened KL, forward vs reverse KL;
(2) a 13-gram overlap contamination check between a benchmark and training text.
"""
import math
def softmax(z, T=1.0):
m = max(z); e = [math.exp((v - m) / T) for v in z]; s = sum(e)
return [v / s for v in e]
def kl(p, q):
return sum(pi * math.log(pi / qi) for pi, qi in zip(p, q) if pi > 1e-12)
# (1a) Temperature reveals "dark knowledge": the teacher's ranking of wrong answers.
teacher_logits = [6.0, 4.5, 1.0, -2.0] # classes: cat, lynx, dog, truck
for T in (1, 2, 4):
print(f"teacher probs at T={T}: {[round(p, 3) for p in softmax(teacher_logits, T)]}")
student_logits = [5.0, 1.0, 3.5, -1.0] # right top-1, wrong idea of what is similar
for T in (1, 4):
loss = (T ** 2) * kl(softmax(teacher_logits, T), softmax(student_logits, T))
print(f"distill loss T^2*KL(teacher||student) at T={T}: {loss:.3f} (T^2 keeps gradient scale comparable)")
print("higher T exposes the student's wrong similarity ranking (dog above lynx) that T=1 mostly hides")
# (1b) Forward vs reverse KL on a bimodal teacher, student restricted to ONE peak shape.
teacher = [0.45, 0.05, 0.0001, 0.05, 0.45] # two good answers (indices 0 and 4)
cands = {"cover both (spread)": [0.25, 0.15, 0.2, 0.15, 0.25],
"seek one mode": [0.90, 0.05, 0.03, 0.01, 0.01]}
print("\nstudent choice forward KL(t||s) reverse KL(s||t)")
for name, s in cands.items():
print(f"{name:<24}{kl(teacher, s):>16.3f}{kl(s, teacher):>18.3f}")
print("forward KL (classic logit distillation) prefers covering both modes; reverse KL (on-policy")
print("distillation: MiniLLM uses reverse KL, GKD trains on student samples with reverse KL or JSD)")
print("prefers committing to one good mode.")
# (2) Contamination: does a benchmark item appear (near-)verbatim in training data?
def ngrams(text, n=13):
toks = text.lower().split()
return {tuple(toks[i:i + n]) for i in range(len(toks) - n + 1)}
train_corpus = ("forum post: here is the answer to the classic puzzle a bat and a ball cost one dollar "
"and ten cents in total the bat costs one dollar more than the ball how much does the ball cost "
"answer five cents")
bench = {
"q1 (leaked)": "a bat and a ball cost one dollar and ten cents in total the bat costs one dollar more than the ball how much does the ball cost",
"q2 (paraphrased)": "together a racket and a shuttlecock are priced at 1.10 and the racket is 1.00 more than the shuttlecock what is the shuttlecock price",
}
train_grams = ngrams(train_corpus)
for name, q in bench.items():
g = ngrams(q)
hit = len(g & train_grams) / max(1, len(g))
print(f"{name:<18} 13-gram overlap={hit:.0%} -> {'FLAG: report separately or drop' if hit > 0.5 else 'passes n-gram check'}")
print("Failure path: q2 passes the n-gram check but is the same problem; paraphrase leaks need embedding")
print("similarity search or fresh, post-cutoff items (dynamic benchmarks) to catch.")Output:
teacher probs at T=1: [0.813, 0.181, 0.005, 0.0]
teacher probs at T=2: [0.636, 0.3, 0.052, 0.012]
teacher probs at T=4: [0.474, 0.326, 0.136, 0.064]
distill loss T^2*KL(teacher||student) at T=1: 0.445 (T^2 keeps gradient scale comparable)
distill loss T^2*KL(teacher||student) at T=4: 2.078 (T^2 keeps gradient scale comparable)
higher T exposes the student's wrong similarity ranking (dog above lynx) that T=1 mostly hides
student choice forward KL(t||s) reverse KL(s||t)
cover both (spread) 0.418 1.556
seek one mode 1.481 0.741
forward KL (classic logit distillation) prefers covering both modes; reverse KL (on-policy
distillation: MiniLLM uses reverse KL, GKD trains on student samples with reverse KL or JSD)
prefers committing to one good mode.
q1 (leaked) 13-gram overlap=100% -> FLAG: report separately or drop
q2 (paraphrased) 13-gram overlap=0% -> passes n-gram check
Failure path: q2 passes the n-gram check but is the same problem; paraphrase leaks need embedding
similarity search or fresh, post-cutoff items (dynamic benchmarks) to catch.Evals: what each kind is for
| Eval type | Measures | Strength | Weakness | Example |
|---|---|---|---|---|
| Static benchmarks (multiple choice, exact match) | Capability on fixed tasks | Cheap, reproducible, comparable | Saturate, contaminate, poorly match your product | MMLU-style, GSM8K-style, HumanEval-style |
| Execution-based | Correctness via running code or checking answers | Objective | Only for checkable tasks | Unit tests, SQL result checks |
| LLM-as-judge (pairwise or rubric) | Open-ended quality, instruction following | Scales, close to human preference when calibrated | Position, length, self-preference bias | MT-Bench-style, Arena-style comparisons |
| Human evaluation | Ground truth for preferences | Trusted | Slow, costly, noisy | Side-by-side ratings |
| Product regression set | Your real tasks and past incidents | Directly relevant | Needs curation and refresh | Golden prompts per feature and slice |
| Safety and red-team suites | Refusals, jailbreak resistance, harmful output | Required for release | Adversaries adapt | Policy-specific prompts |
| Online A/B and canary | Real user outcomes | The final truth | Slow, risky | Thumbs, retention, task success |
LLM-as-judge done right. Use a strong judge with a written rubric; prefer pairwise comparison to absolute scores; swap positions and count disagreements as ties; control for length; avoid judging a model with itself; calibrate against a few hundred human labels and report judge-human agreement. The MT-Bench paper measured GPT-4-as-judge agreement with humans at a level similar to human-human agreement, along with the biases above.
Regression gates
A gate is a release rule, written before you look at the numbers:
- Overall improvement must be statistically real: a paired comparison (same items for both models) with a bootstrap confidence interval that excludes zero.
- No protected slice may regress beyond a tolerance: safety refusals, each language, each customer-critical task family. Averages hide regressions.
- Hard floors on safety and format-validity metrics regardless of averages.
- Judge sanity checks: position-swap agreement above a threshold, judge-human agreement measured.
// Eval regression gate for a candidate model vs the current production model.
// (1) Paired bootstrap on per-item scores -> confidence interval for the delta.
// (2) Per-slice gates: no protected slice may regress beyond a tolerance, even if the average improves.
// (3) LLM-as-judge position-bias check: judge each pair twice with A/B swapped; count flips.
function lcg(seed: number): () => number {
let s = seed >>> 0;
return () => ((s = (Math.imul(s, 1664525) + 1013904223) >>> 0) / 2 ** 32);
}
const rnd = lcg(2026);
type Item = { slice: string; prod: number; cand: number };
const slices: Array<[string, number, number, number]> = [
// name, items, prod pass rate, candidate pass rate (synthetic)
["coding", 400, 0.62, 0.70],
["math", 300, 0.55, 0.60],
["safety-refusals", 200, 0.97, 0.91], // the regression hiding under a better average
["multilingual", 100, 0.70, 0.71],
];
const items: Item[] = [];
for (const [slice, n, p, c] of slices) {
for (let i = 0; i < n; i++) {
const u = rnd(); // shared noise makes scores paired (same item difficulty for both models)
items.push({ slice, prod: u < p ? 1 : 0, cand: u < c ? 1 : 0 });
}
}
function bootstrapDelta(xs: Item[], iters = 2000): [number, number, number] {
const deltas: number[] = [];
for (let b = 0; b < iters; b++) {
let d = 0;
for (let i = 0; i < xs.length; i++) {
const it = xs[Math.floor(rnd() * xs.length)];
d += it.cand - it.prod;
}
deltas.push(d / xs.length);
}
deltas.sort((a, b) => a - b);
const mean = xs.reduce((s, it) => s + it.cand - it.prod, 0) / xs.length;
return [mean, deltas[Math.floor(0.025 * iters)], deltas[Math.floor(0.975 * iters)]];
}
const fmt = (x: number) => (x >= 0 ? "+" : "") + (100 * x).toFixed(1) + "pp";
const [m, lo, hi] = bootstrapDelta(items);
console.log(`overall delta ${fmt(m)} 95% CI [${fmt(lo)}, ${fmt(hi)}]`);
let pass = lo > 0;
const tolerance = -0.02; // a protected slice may not drop more than 2 points (upper CI bound)
for (const [slice] of slices) {
const [sm, slo, shi] = bootstrapDelta(items.filter((it) => it.slice === slice));
const bad = shi < 0 || sm < tolerance;
if (bad) pass = false;
console.log(` ${slice.padEnd(16)} ${fmt(sm).padStart(8)} CI [${fmt(slo)}, ${fmt(shi)}]${bad ? " <- REGRESSION" : ""}`);
}
// Judge position bias: a biased judge prefers whichever answer is shown first 30% of the time.
let flips = 0;
const N = 500;
for (let i = 0; i < N; i++) {
const trueBetterIsA = rnd() < 0.5;
const judge = (firstIsA: boolean) => (rnd() < 0.3 ? firstIsA : trueBetterIsA); // returns "A wins?"
if (judge(true) !== judge(false)) flips++;
}
console.log(`\njudge verdict flipped when A/B order swapped on ${((100 * flips) / N).toFixed(0)}% of pairs -> average both orders, count flips as ties`);
console.log(`GATE: ${pass ? "PASS" : "FAIL"} (overall CI must exclude 0 AND no protected slice may regress)`);Output:
overall delta +4.3pp 95% CI [+2.8pp, +5.9pp]
coding +8.3pp CI [+5.5pp, +11.0pp]
math +6.7pp CI [+4.0pp, +9.7pp]
safety-refusals -6.5pp CI [-10.0pp, -3.5pp] <- REGRESSION
multilingual +3.0pp CI [+0.0pp, +7.0pp]
judge verdict flipped when A/B order swapped on 30% of pairs -> average both orders, count flips as ties
GATE: FAIL (overall CI must exclude 0 AND no protected slice may regress)Expectedoverall delta +4.3pp 95% CI [+2.8pp, +5.9pp] coding +8.3pp CI [+5.5pp, +11.0pp] math +6.7pp CI [+4.0pp, +9.7pp] safety-refusals -6.5pp CI [-10.0pp, -3.5pp] <- REGRESSION multilingual +3.0pp CI [+0.0pp, +7.0pp] judge verdict flipped when A/B order swapped on 30% of pairs -> average both orders, count flips as ties GATE: FAIL (overall CI must exclude 0 AND no protected slice may regress)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Contamination
If test items, or close paraphrases, appeared in pretraining or fine-tuning data, scores measure memory, not capability. Detection and defense, in increasing strength:
- n-gram overlap (the GPT-3 paper used 13-grams) between eval items and training corpora: cheap, misses paraphrases.
- Embedding similarity search for near-duplicates and paraphrases.
- Behavioural probes: the model completes a benchmark question verbatim, or performance drops sharply on perturbed versions.
- Fresh, post-cutoff items (dynamic benchmarks such as LiveBench), private held-out sets that never leave your infrastructure, and canary strings in eval files.
The release pipeline
Diagram 1: from candidate checkpoint to production, with the failure path.
Decisions
- 1
Step 1: candidate checkpoint (SFT, DPO, GRPO or distilled student)
- nextStep 2: decontamination check of training data vs eval sets
- 2
Step 2: decontamination check of training data vs eval sets
- nextStep 3: offline benchmarks plus execution-based tests
- 3
Step 3: offline benchmarks plus execution-based tests
- nextStep 4: LLM-as-judge pairwise vs production, both orders, length controlled
- 4
Step 4: LLM-as-judge pairwise vs production, both orders, length controlled
- nextStep 5: regression gate - paired bootstrap CI, per-slice tolerances, safety floors
- 5
Step 5: regression gate - paired bootstrap CI, per-slice tolerances, safety floors
- nextStep 6: gate passes?
- ?
Step 6: gate passes?
- nextStep 7: canary to a small traffic share, watch online metrics
- nextFailure path: average up but a protected slice regressed, or judge flips on order swap
- 7
Step 7: canary to a small traffic share, watch online metrics
- nextStep 8: full rollout, keep previous model warm for rollback
- 8
Step 8: full rollout, keep previous model warm for rollback
- 9
Failure path: average up but a protected slice regressed, or judge flips on order swap
- nextBlock release, trace the regression to its data or stage, add the failing cases to the regression set
- 10
Block release, trace the regression to its data or stage, add the failing cases to the regression set
- nextStep 1: candidate checkpoint (SFT, DPO, GRPO or distilled student)
Lesson map
Distillation & LLM Evals - Teacher-Student Training, Benchmarks, LLM-as-Judge, Regression Gates & Contamination
Sequence-level vs logit-level vs on-policy distillation and reasoning-trace distillation; benchmarks, LLM-as-judge biases, regression gates with per-slice thresholds and CIs, contamination checks; runnable distill/contamination + regression gate.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: candidate checkpoint (SFT, DPO, GRPO or distilled student)"] s2["Step 2: decontamination check of training data vs eval sets"] s3["Step 3: offline benchmarks plus execution-based tests"] s4["Step 4: LLM-as-judge pairwise vs production, both orders, length controlled"] s5["Step 5: regression gate - paired bootstrap CI, per-slice tolerances, safety floors"] s6["Step 6: gate passes?"] s7["Step 7: canary to a small traffic share, watch online metrics"] s8["Step 8: full rollout, keep previous model warm for rollback"] f1["Failure path: average up but a protected slice regressed, or judge flips on order swap"] f2["Block release, trace the regression to its data or stage, add the failing cases to the regression set"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s4 -->|continues| s5 s5 -->|continues| s6 s6 -->|continues| s7 s7 -->|continues| s8 s6 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s1
Decision chart: distill, and how to evaluate?
Decisions
- 1
D1
- nextDo not distill, ship the larger model
- nextStep 2: white-box teacher with the same tokenizer?
- 2
Do not distill, ship the larger model
- ?
Step 2: white-box teacher with the same tokenizer?
- nextSequence-level distillation: generate, filter, SFT (check licence terms)
- nextStep 3: can you afford on-policy sampling?
- 4
Sequence-level distillation: generate, filter, SFT (check licence terms)
- nextStep 4: is the task checkable by a program?
- ?
Step 3: can you afford on-policy sampling?
- nextOn-policy distillation with reverse KL or JSD (GKD, MiniLLM style)
- nextLogit distillation with temperature, forward KL
- 6
On-policy distillation with reverse KL or JSD (GKD, MiniLLM style)
- nextStep 4: is the task checkable by a program?
- 7
Logit distillation with temperature, forward KL
- nextStep 4: is the task checkable by a program?
- ?
Step 4: is the task checkable by a program?
- nextExecution-based evals plus regression slices
- nextPairwise LLM-as-judge calibrated on human labels, plus regression slices
- 9
Execution-based evals plus regression slices
- 10
Pairwise LLM-as-judge calibrated on human labels, plus regression slices
What happens if you choose otherwise
- Distill and ship without a gap measurement: users find the cases where the student is worse; measure the teacher-student gap per slice and decide which slices can route to the teacher.
- Single-number eval: a better average hides a safety regression; the release looks like a win until incidents arrive.
- Absolute 1-10 judge scores: drift between judge versions and poor discrimination; pairwise with swaps is more stable.
- Public benchmarks only: contamination and saturation make them a weak signal for your product; a private regression set is what catches real breakage.
Pitfalls
- Using the same model family as teacher, student and judge, then trusting its preferences.
- Evaluating with different decoding settings (temperature, max tokens, chat template) than production.
- Not versioning eval sets, so scores are not comparable across releases.
- Treating small deltas on a few hundred items as real without confidence intervals.
Interview Q&A
Logit distillation vs sequence-level distillation?
Answer
Logit distillation trains on the teacher's full distributions (needs white-box access and the same vocabulary) and carries more signal per token; sequence-level trains on generated text, works through an API, and is simpler but loses uncertainty information.
Why use temperature in distillation?
Answer
A softened teacher distribution exposes the relative probabilities of wrong answers, which encode similarity structure; scale the loss by T squared to keep gradients comparable.
How do you make an LLM judge trustworthy?
Answer
Rubric, pairwise comparisons, position swaps with ties for disagreements, length control, a judge different from the model under test, and measured agreement with human labels.
What belongs in a regression gate?
Answer
A paired statistical test on overall quality with a CI, per-slice no-regression tolerances, hard floors on safety and format validity, and judge sanity checks; rules fixed before results are seen.
How do you detect benchmark contamination?
Answer
n-gram overlap against training data, embedding similarity for paraphrases, perturbation and completion probes, and relying on fresh or private held-out sets.
What is on-policy distillation?
Answer
The teacher scores the student's own samples (MiniLLM, GKD), fixing exposure bias at the cost of online sampling.
Why prefer pairwise judging over absolute 1-10 scores?
Answer
Absolute scores drift between judge versions and discriminate poorly; pairwise with swaps is more stable.
Why scale the distillation loss by T squared?
Answer
To keep gradient scale comparable across temperatures (Hinton et al.).
Check yourself
Take your current eval set and split it into protected slices. Write the gate rules (overall CI, per-slice tolerance, hard floors, judge checks) before running the next candidate.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Agent Reliability — Evals, Guardrails, Tracing & Failure Modes, RAG Failure Modes — Hallucination, Stale Indexes, Evals & Grounding, Speculative Decoding, Algorithm Selection Playbook — Metrics, Baselines & Failure Modes, Load, Chaos & Production Validation — Shadow Traffic, Canaries & Game Days, Feature Flags — Targeting, Experimentation & Kill Switches, LLM Inference in Production - Quantization, Prompt Caching, Multi-LoRA Serving & Capacity Planning.
Go Deeper
- Distilling the Knowledge in a Neural Network (Hinton et al.)
- MiniLLM: On-Policy Distillation of Large Language Models
- On-Policy Distillation of Language Models (GKD)
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Language Models are Few-Shot Learners (13-gram contamination analysis)
- LiveBench: a contamination-limited benchmark
- EleutherAI lm-evaluation-harness
- Stanford HELM