Supervised Fine-Tuning (SFT) - Instruction Data, Chat Templates, Packing & Loss Masking
SFT data formats, chat templates and special tokens, assistant-only loss masking, packing with attention boundaries, truncation; data quality over volume; runnable masking + packing demos; template/masking bug failure paths.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is SFT?
Answer
Next-token training on (prompt, ideal response) examples, formatted with a chat template.
L2
Why mask prompt tokens?
Answer
The model should learn to produce responses, not to write system prompts or user questions.
L3
What is a chat template?
Answer
A deterministic rendering of role-tagged messages into tokens with special markers for roles and turn ends.
L4
Why train the end-of-turn token?
Answer
If the loss never covers it, the model never learns to stop.
L5
Padding vs packing?
Answer
Padding wastes compute on pad tokens; packing fills rows but needs position resets and a block-diagonal mask.
L6
How much data do you need?
Answer
Hundreds to low thousands of high-quality, diverse examples can work for behaviour and format.
L7
Why did the model forget instructions after SFT?
Answer
Narrow data mix, too high a learning rate or too many epochs.
Failure modes
No loss masking
The model learns to write user turns and may role-play the user in production.
Naive packing
Each example attends to unrelated earlier examples in the same row.
Truncated answers
Cutting off the end of assistant answers teaches the model to stop mid-sentence.
Misconceptions
SFT is a good way to inject new knowledge.
It mostly teaches format and style; new knowledge through SFT raises hallucination risk.
More synthetic data is always better.
One teacher prompt at scale collapses diversity.
Any reasonable template works at serving time.
A different template makes every request slightly out of distribution.
Interviewer traps
Adding special tokens without training them.
Resize embeddings and train the new tokens; untrained embeddings are random vectors.
Reading training loss as quality.
Watch held-out loss and evals.
Design scenario
Same prompt for every reader.
Requirements
Correct format, reliable stopping, tool-call syntax matching the serving parser, and no forgetting of general instructions.
Failure assumptions
- Logs contain PII.
- Some examples exceed the sequence length.
- The serving stack uses a different template version.
Constraints
- One training run per week.
- Mixed-length data from 50 to 4,000 tokens.
Prompt
Fine-tune an open model into a terse SQL assistant from a mix of human demos and filtered production logs.
API
Which message format and chat template do training and serving share?
Data
How are examples deduplicated, filtered, masked and packed?
Architecture
Which batching strategy, learning rate and data mix does the run use?
Overview
Supervised fine-tuning (SFT), also called instruction tuning, trains a pretrained model on (prompt, ideal response) examples with the same next-token loss used in pretraining, with two crucial differences: the text is formatted with a chat template that marks roles and turn boundaries with special tokens, and the loss is masked so the model only learns to produce assistant tokens, not to reproduce system prompts or user questions. Most of the quality of SFT comes from the data (diversity, correctness, consistent style) rather than from volume; most of the bugs come from plumbing (templates, masking, packing, truncation). This page covers both.
What SFT data looks like, compared
| Format | Shape | Good for | Watch out for |
|---|---|---|---|
| Instruction / response (Alpaca-style) | {instruction, input, output} | Single-turn tasks | Has no notion of multi-turn context or system prompts |
| Chat messages | [{role: system/user/assistant/tool, content}] | Assistants, multi-turn, tool use | Must be rendered with the exact template the model will be served with |
| Prompt / completion | Raw prompt + completion strings | Classification-like tasks, completion models | Easy to forget an end-of-sequence token, so the model never stops |
| Tool-call traces | Messages with assistant tool calls and tool results | Agents, function calling | Train on the assistant's call arguments, not on tool outputs it cannot control |
Where the data comes from matters as much as the format:
- Human-written demonstrations: highest trust, slowest and most expensive, inconsistent style across annotators.
- Synthetic data from a stronger model (distillation by generation, Self-Instruct style): cheap and consistent, inherits the teacher's errors and licence terms, can collapse diversity.
- Filtered production logs: realistic distribution, needs consent, PII scrubbing and quality filtering.
The LIMA result (about 1,000 carefully curated examples producing a strong assistant) is the usual citation for "quality over quantity": SFT mostly teaches format and style on top of knowledge the base model already has. Teaching genuinely new knowledge through SFT is unreliable and increases hallucination risk, because the model learns to answer confidently about things its weights do not support.
Pad every row or pack examples together?
Prefer
Packing with attention isolation
Concatenate examples into full rows, reset positions, use a block-diagonal mask.
- Utilisation rose from 19.8% to 99.8%.
- About 5.0x fewer rows for the same 4,000 examples.
- With the block-diagonal mask the last token sees 0 tokens from other examples.
Alternative
Pad to the longest length
One example per row, padded.
- Simplest to implement.
- Most of each row can be padding, so compute is wasted.
- Length bucketing recovers much of the waste without attention tricks.
One SFT run
Diagram 1 condensed: data to checkpoint, with the common failure path.
- 1
Curate data
Dedupe, filter quality and scrub PII. - 2
Render with the chat template
The same template the server will use. - 3
Mask and batch
Loss on assistant tokens only; pack with position resets. - 4
Train briefly
Low learning rate, 1-3 epochs, some general data. - 5
Template or mask bug
Quality drops silently until a side-by-side eval.
Chat templates: the contract between training and serving
A chat template renders a list of messages into one token sequence. Llama 3 uses headers like <|start_header_id|>assistant<|end_header_id|> and <|eot_id|>; ChatML-style templates (used by Qwen and others) use <|im_start|>assistant and <|im_end|>. Hugging Face tokenizers ship the template as a Jinja string so apply_chat_template renders the same way in training and serving.
Rules that prevent silent quality loss:
- Train and serve with the identical template. A server that formats prompts differently makes every request slightly out of distribution.
- Train the end-of-turn token. If the loss never covers
<|eot_id|>or<|im_end|>, the model never learns to stop. - Add the generation prompt at inference (the
assistantheader) so the model starts answering instead of predicting a user turn. - When you add new special tokens, resize embeddings and train them; untrained new embeddings are random vectors.
Loss masking
Labels for system, user and tool-result tokens are set to an ignore index (-100 in PyTorch cross-entropy), so gradients come only from assistant content. Alternatives and when they make sense:
| Strategy | Loss on | When to use | Risk |
|---|---|---|---|
| Assistant-only (default for chat) | All assistant turns + end-of-turn | Assistants, multi-turn | None significant |
| Completion-only, last turn | Final assistant turn only | When earlier turns are synthetic context you do not want to imitate | Wastes signal from good earlier turns |
| Full sequence | Everything | Continued pretraining on domain text | Model learns to write user turns and system prompts |
"""SFT data plumbing: chat template -> tokens -> labels with loss masking.
Real trainers use the model's own tokenizer and its chat template (for example a
Jinja template shipped with the tokenizer). Here a whitespace "tokenizer" keeps the
mechanics visible. The key rule: the loss is computed only on tokens the model
should learn to PRODUCE (assistant turns + end-of-turn marker), never on the
system/user text it is merely conditioned on. Ignored positions get label -100,
the convention PyTorch's CrossEntropyLoss(ignore_index=-100) uses.
"""
import math, random
IGNORE = -100
def render_chatml(messages):
"""Render to a ChatML-style string and remember which spans are assistant output."""
pieces = [] # (text, trainable)
for m in messages:
pieces.append((f"<|im_start|> {m['role']}", False))
pieces.append((m["content"], m["role"] == "assistant"))
# The end-of-turn token after an assistant reply IS trained: the model must learn to stop.
pieces.append(("<|im_end|>", m["role"] == "assistant"))
return pieces
def tokenize(pieces, vocab):
ids, labels = [], []
for text, trainable in pieces:
for tok in text.split():
idx = vocab.setdefault(tok, len(vocab))
ids.append(idx)
labels.append(idx if trainable else IGNORE)
return ids, labels
convo = [
{"role": "system", "content": "You are a terse SQL assistant ."},
{"role": "user", "content": "count users created today"},
{"role": "assistant", "content": "SELECT count(*) FROM users WHERE created_at::date = current_date ;"},
{"role": "user", "content": "only active ones"},
{"role": "assistant", "content": "SELECT count(*) FROM users WHERE active AND created_at::date = current_date ;"},
]
vocab = {}
ids, labels = tokenize(render_chatml(convo), vocab)
inv = {v: k for k, v in vocab.items()}
# Compact view: trained tokens are wrapped in [brackets], ignored ones are plain.
view = " ".join(f"[{inv[i]}]" if l != IGNORE else inv[i] for i, l in zip(ids, labels))
for line in view.replace("<|im_start|>", "\n<|im_start|>").strip().splitlines():
print(" ", line.rstrip())
trained = sum(l != IGNORE for l in labels)
print(f"\n{trained}/{len(labels)} tokens contribute to the loss ({trained/len(labels):.0%})")
# Shifted next-token loss on fake logits: position t predicts token t+1.
random.seed(0)
V = len(vocab)
def xent(logits, target):
m = max(logits); lse = m + math.log(sum(math.exp(x - m) for x in logits))
return lse - logits[target]
logits = [[random.gauss(0, 1) for _ in range(V)] for _ in ids]
def mean_loss(mask):
tot, n = 0.0, 0
for t in range(len(ids) - 1):
tgt = labels[t + 1] if mask else ids[t + 1]
if tgt == IGNORE:
continue
tot += xent(logits[t], tgt); n += 1
return tot / n, n
for mask in (True, False):
loss, n = mean_loss(mask)
print(f"masked={mask!s:<5} loss over {n:2d} positions = {loss:.3f}")
print("\nFailure path: with masked=False the model is also trained to WRITE system prompts and user")
print("questions, wasting capacity and teaching it to role-play the user. A wrong template at")
print("inference (different special tokens than training) silently degrades quality.")Output:
<|im_start|> system You are a terse SQL assistant . <|im_end|>
<|im_start|> user count users created today <|im_end|>
<|im_start|> assistant [SELECT] [count(*)] [FROM] [users] [WHERE] [created_at::date] [=] [current_date] [;] [<|im_end|>]
<|im_start|> user only active ones <|im_end|>
<|im_start|> assistant [SELECT] [count(*)] [FROM] [users] [WHERE] [active] [AND] [created_at::date] [=] [current_date] [;] [<|im_end|>]
22/49 tokens contribute to the loss (45%)
masked=True loss over 22 positions = 3.729
masked=False loss over 48 positions = 3.828
Failure path: with masked=False the model is also trained to WRITE system prompts and user
questions, wasting capacity and teaching it to role-play the user. A wrong template at
inference (different special tokens than training) silently degrades quality.Batching: padding, bucketing and packing
Sequences vary wildly in length. Three ways to batch them, from most to least waste:
- Pad to max length: simplest, but most of each row can be padding. Compute on pad tokens is pure waste.
- Dynamic padding with length bucketing: sort or bucket by length and pad to the longest in the batch. Most of the gain, no attention tricks needed.
- Packing: concatenate several examples into one full-length row. Near 100 percent utilisation, but you must reset position ids per example and use a block-diagonal attention mask (FlashAttention's variable-length kernels support this), otherwise each example attends to unrelated earlier examples in the same row.
// Padding vs packing for SFT batches, with the cross-contamination detail.
// Padding: every example sits alone in a row of length L; the rest is pad (wasted compute).
// Packing: concatenate several examples into one row of length L (first-fit decreasing here).
// Naive packing lets example B attend to example A in the same row; correct packing resets
// position ids and uses a block-diagonal ("document") attention mask so examples stay isolated.
function lcg(seed: number): () => number {
let s = seed >>> 0;
return () => ((s = (Math.imul(s, 1664525) + 1013904223) >>> 0) / 2 ** 32);
}
const L = 2048; // max sequence length per row
const rnd = lcg(7);
// SFT data is usually long-tailed: many short chats, a few long ones.
const lengths = Array.from({ length: 4000 }, () => Math.min(L, Math.floor(60 + 1400 * rnd() ** 3)));
const tokens = lengths.reduce((a, b) => a + b, 0);
// Padding: one example per row.
const padRows = lengths.length;
const padUtil = tokens / (padRows * L);
// Packing: first-fit decreasing bin packing.
const bins: number[][] = [];
const free: number[] = [];
for (const len of [...lengths].sort((a, b) => b - a)) {
let i = free.findIndex((f) => f >= len);
if (i === -1) { bins.push([]); free.push(L); i = bins.length - 1; }
bins[i].push(len);
free[i] -= len;
}
const packUtil = tokens / (bins.length * L);
console.log(`examples=${lengths.length} real tokens=${tokens}`);
console.log(`padding : rows=${padRows} utilisation=${(padUtil * 100).toFixed(1)}%`);
console.log(`packing : rows=${bins.length} utilisation=${(packUtil * 100).toFixed(1)}% -> ~${(padRows / bins.length).toFixed(1)}x fewer rows`);
// Position ids and attention isolation inside one packed row.
const row = bins[bins.length - 1];
const positionIds: number[] = [];
const docIds: number[] = [];
row.forEach((len, d) => { for (let p = 0; p < len; p++) { positionIds.push(p); docIds.push(d); } });
const canAttend = (q: number, k: number, isolated: boolean) => k <= q && (!isolated || docIds[q] === docIds[k]);
const q = positionIds.length - 1; // last token of the last example in this row
let leakNaive = 0, leakIsolated = 0;
for (let k = 0; k <= q; k++) {
if (docIds[k] !== docIds[q]) {
if (canAttend(q, k, false)) leakNaive++;
if (canAttend(q, k, true)) leakIsolated++;
}
}
console.log(`\nlast packed row holds ${row.length} examples; position ids restart at 0 for each: [${positionIds.slice(0, 3).join(",")}, ... ${positionIds.slice(row[0] - 1, row[0] + 2).join(",")} ...]`);
console.log(`tokens from OTHER examples the last token can see: naive causal mask=${leakNaive}, block-diagonal mask=${leakIsolated}`);
console.log("Failure path: naive packing trains on cross-example context; losses look fine but the model learns spurious dependencies.");Output:
examples=4000 real tokens=1620521
padding : rows=4000 utilisation=19.8%
packing : rows=793 utilisation=99.8% -> ~5.0x fewer rows
last packed row holds 27 examples; position ids restart at 0 for each: [0,1,2, ... 59,0,1 ...]
tokens from OTHER examples the last token can see: naive causal mask=1560, block-diagonal mask=0
Failure path: naive packing trains on cross-example context; losses look fine but the model learns spurious dependencies.Expectedexamples=4000 real tokens=1620521 padding : rows=4000 utilisation=19.8% packing : rows=793 utilisation=99.8% -> ~5.0x fewer rows last packed row holds 27 examples; position ids restart at 0 for each: [0,1,2, ... 59,0,1 ...] tokens from OTHER examples the last token can see: naive causal mask=1560, block-diagonal mask=0 Failure path: naive packing trains on cross-example context; losses look fine but the model learns spurious dependencies.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The SFT run, step by step
Diagram 1: one SFT run with the failure path that shows up most in practice.
Decisions
- 1
Step 1: collect demos (human, synthetic from a teacher, filtered logs)
- nextStep 2: dedupe, filter, decontaminate against eval sets
- 2
Step 2: dedupe, filter, decontaminate against eval sets
- nextStep 3: render with the model's chat template
- 3
Step 3: render with the model's chat template
- nextStep 4: tokenize, build labels, mask non-assistant tokens with -100
- 4
Step 4: tokenize, build labels, mask non-assistant tokens with -100
- nextStep 5: pack or bucket into batches, reset position ids per example
- 5
Step 5: pack or bucket into batches, reset position ids per example
- nextStep 6: train 1-3 epochs, small learning rate, full FT or LoRA
- 6
Step 6: train 1-3 epochs, small learning rate, full FT or LoRA
- nextStep 7: held-out evals and a general-capability regression set
- ?
Step 7: held-out evals and a general-capability regression set
- nextStep 8: hand to preference tuning or deploy
- nextFailure path: loss looks great but answers never stop, or quality drops elsewhere
- 8
Step 8: hand to preference tuning or deploy
- 9
Failure path: loss looks great but answers never stop, or quality drops elsewhere
- nextCheck end-of-turn labels, template mismatch, over-training, missing general data in the mix
- 10
Check end-of-turn labels, template mismatch, over-training, missing general data in the mix
- nextStep 3: render with the model's chat template
Lesson map
Supervised Fine-Tuning (SFT) - Instruction Data, Chat Templates, Packing & Loss Masking
SFT data formats, chat templates and special tokens, assistant-only loss masking, packing with attention boundaries, truncation; data quality over volume; runnable masking + packing demos; template/masking bug failure paths.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: collect demos (human, synthetic from a teacher, filtered logs)"] s2["Step 2: dedupe, filter, decontaminate against eval sets"] s3["Step 3: render with the model's chat template"] s4["Step 4: tokenize, build labels, mask non-assistant tokens with -100"] s5["Step 5: pack or bucket into batches, reset position ids per example"] s6["Step 6: train 1-3 epochs, small learning rate, full FT or LoRA"] s7["Step 7: held-out evals and a general-capability regression set"] s8["Step 8: hand to preference tuning or deploy"] f1["Failure path: loss looks great but answers never stop, or quality drops elsewhere"] f2["Check end-of-turn labels, template mismatch, over-training, missing general data in the mix"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s4 -->|continues| s5 s5 -->|continues| s6 s6 -->|continues| s7 s7 -->|continues| s8 s7 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s3
Hyperparameters that matter
- Learning rate: much lower than pretraining for full fine-tuning (on the order of 1e-5); LoRA tolerates higher (on the order of 1e-4). Too high causes forgetting.
- Epochs: 1-3 is typical. More epochs on small data memorise it; watch held-out loss and evals, not training loss.
- Data mix: add a slice of general instruction data alongside your domain data to reduce catastrophic forgetting of skills you are not training.
- Sequence length: truncation that cuts off the end of assistant answers teaches the model to stop mid-sentence. Filter or split long examples instead.
- NEFTune-style embedding noise and similar tricks are optional; data quality dominates.
Decision chart: how should this SFT run be set up?
Decisions
- 1
D1
- nextChat messages format plus the model's own template
- nextPrompt and completion format, remember EOS
- 2
Chat messages format plus the model's own template
- nextStep 2: are earlier turns trustworthy examples?
- 3
Prompt and completion format, remember EOS
- nextStep 3: very uneven sequence lengths?
- ?
Step 2: are earlier turns trustworthy examples?
- nextLoss on all assistant turns
- nextLoss on the final assistant turn only
- ?
Step 3: very uneven sequence lengths?
- nextPacking with position reset and block-diagonal mask
- nextDynamic padding with length buckets
- 6
Loss on all assistant turns
- nextStep 3: very uneven sequence lengths?
- 7
Loss on the final assistant turn only
- nextStep 3: very uneven sequence lengths?
- 8
Packing with position reset and block-diagonal mask
- nextStep 4: teaching new facts rather than behaviour?
- 9
Dynamic padding with length buckets
- nextStep 4: teaching new facts rather than behaviour?
- ?
Step 4: teaching new facts rather than behaviour?
- nextPrefer RAG or continued pretraining, SFT alone hallucinates
- nextProceed, keep a general-data slice to limit forgetting
- 11
Prefer RAG or continued pretraining, SFT alone hallucinates
- 12
Proceed, keep a general-data slice to limit forgetting
What happens if you choose otherwise
- No loss masking: the model learns to write user questions and system prompts; in production it may continue past its turn and role-play the user.
- Different template in serving: answers are subtly worse everywhere, and nobody notices until a side-by-side eval.
- Naive packing without isolation: training loss looks normal, but the model learns spurious cross-example dependencies.
- Huge synthetic dataset from one teacher prompt: consistent style, collapsed diversity; the model gives the same shaped answer to every question.
- SFT to inject private knowledge: confident wrong answers on the edges of that knowledge; RAG is usually the safer tool.
Pitfalls
- Duplicates and near-duplicates in the dataset get effectively up-weighted; dedupe before training.
- Leaking evaluation prompts into SFT data, often through synthetic generation seeded from benchmarks.
- Mixing datasets with conflicting styles (some terse, some verbose) produces an inconsistent assistant.
- Forgetting to train tool-call formatting exactly as the serving parser expects; one schema mismatch breaks every call.
Interview Q&A
Why mask the prompt tokens in SFT?
Answer
The model is conditioned on the prompt but should only learn to produce the response. Training on prompts wastes capacity, biases the model toward writing user-like text, and can make it continue into the next user turn.
What is a chat template and why does it matter?
Answer
A deterministic rendering of role-tagged messages into tokens with special markers for roles and turn ends. Training and serving must use the same template; otherwise every request is slightly out of distribution, and missing end-of-turn training means the model never stops.
Padding vs packing?
Answer
Padding wastes compute on pad tokens; packing fills rows with multiple examples for near-full utilisation, but requires resetting position ids and a block-diagonal attention mask to keep examples independent.
How much SFT data do you need?
Answer
Less than people expect for behaviour and format (hundreds to low thousands of high-quality, diverse examples can work), more for new skills. Quality, diversity and consistency matter more than count.
Your SFT model stopped following instructions it used to follow. Why?
Answer
Catastrophic forgetting from a narrow data mix or too high a learning rate or too many epochs. Mix in general instruction data, lower the learning rate, train fewer steps, or use LoRA to limit drift.
What does the ignore index do in loss masking?
Answer
Labels set to -100 in PyTorch cross-entropy contribute no gradient, so system, user and tool-result tokens are conditioned on but not learned.
Why should tool-call traces train on the assistant's call arguments but not tool outputs?
Answer
The model controls its call arguments but cannot control what tools return.
What does the LIMA result suggest about SFT?
Answer
About 1,000 carefully curated examples produced a strong assistant: SFT mostly teaches format and style on knowledge the base model already has.
Check yourself
Render one multi-turn example from your data with your model's chat template, mark which tokens get loss, and check that the end-of-turn token is included and that no answer is truncated.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Tool/Function Calling vs Structured Outputs, Grammars & CFG Constrained Decoding, Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism.