Speculative Decoding
When decode is memory-bandwidth bound, speculative decoding spends spare FLOPs to cut sequential target steps. A draft proposes several tokens and the target verifies them in one forward. Acceptance rate decides whether you win or waste the step.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Is the draft earning its FLOPs?
Prefer
Verify several tokens while memory is the bottleneck
One target forward scores a draft prefix in parallel. When acceptance is high, you amortize the KV read across multiple output tokens.
- Decode-bound serving with spare FLOPs.
- A draft that matches the domain and the temperature.
- Moderate k. Longer drafts are not free.
Alternative
Always draft eight tokens
Low acceptance pays draft cost plus a verify that keeps one token. Throughput can fall below vanilla decode. Prefill-heavy batches lose the FLOPs they needed for time-to-first-token.
- Domain shift between draft and target.
- Tiny batches where overhead dominates.
- A sampler that does not match the target distribution.
Overview
Decode streams KV and does little math per new token. Speculative decoding uses leftover compute: a cheap draft proposes k tokens ahead, and the target model scores them in one parallel forward. You keep the longest prefix the target agrees with, then repeat. Interviewers want the acceptance math and the failure cases, not a Medusa training recipe.
This page is not a grammar or JSON Schema lesson. Constrained decoding stays in Structured Outputs / Constrained Decoding and SGLang — RadixAttention, Continuous Batching & Structured Generation.
Draft plus verify
- A draft model, or extra heads on the target, proposes
ktokens. - The target runs one forward that scores those draft tokens together (and often predicts one more).
- Accept the longest prefix consistent with the target sampling policy.
- On the first rejection, keep the corrected token from the target, discard the rest, and draft again.
Exact string match under greedy decoding is the toy below. Real speculative sampling uses a correction so the output distribution matches the target. Say which policy you implemented. A buggy accept/reject biases the text.
Acceptance rate
Let α be the chance a drafted token is accepted, and k the number proposed. Higher α means fewer target steps per output token. A standard interview form, for a constant α and k drafted tokens, is that one target verification yields about (1 - α^(k+1)) / (1 - α) tokens when you also keep the corrected token on rejection. If α is 1 the limit is k + 1.
If α is low, you pay the draft and the verify for almost no extra tokens. Throughput drops below vanilla decode. α moves with draft quality, temperature, domain shift, and how far ahead you draft. Do not quote a speedup without an acceptance number from your traffic.
Medusa, EAGLE, tree attention
High level only:
- Medusa-style. Extra decoding heads on the target predict future tokens, so you draft without a second full model.
- EAGLE-style. Draft features come from the target’s hidden states, with iterative draft heads.
- Tree attention. Verify several candidate branches in one target forward using a packed tree mask, instead of a single chain.
All three parallelize verification of speculative tokens. They differ in where the draft comes from and how many candidates you score. Training those heads, and the feature-level loss, is out of scope.
| Draft source | Gain | Cost |
|---|---|---|
| Separate small model | Flexible pairing | Extra weights, sync, poor α if the draft mismatches |
| Medusa or EAGLE heads | Colocated with the target | Heads to maintain, architecture-specific |
| Tree verify | More candidates per step | Denser masks, more transient KV |
| Vanilla decode | Predictable | Leaves decode bandwidth unused |
When it helps and when it hurts
Helps when serving is decode-bound, the draft is good, k is moderate, and the GPU has spare FLOPs while HBM waits on KV.
Hurts when the workload is prefill-dominated, the draft is off-domain, k is large, the GPU is already compute-saturated, or the batch is so small that launch overhead dominates.
A prefix cache still helps the prompt. Speculation mainly accelerates the decode tail. It does not skip a cold prefill. That split is Prefill vs Decode & KV Cache Mechanics.
Continuous batching and KV
Speculative steps change how many tokens a sequence advances in one iteration. The scheduler in PagedAttention & Continuous Batching must account for a variable advance and for draft tokens that share a verify batch.
Paging still applies. Accepted tokens append KV. Rejected drafts must not stay in the cache. If you wrote them speculatively, roll them back. Leaving rejected KV in the block is silent corruption.
Flow
- 1
1 Draft model or extra heads
- next2 Candidate tokens ahead
- 2
2 Candidate tokens ahead
- next3 Target forward in parallel
- 3
3 Target forward in parallel
- next4 Accept the matching prefix
- 4
4 Accept the matching prefix
- next5 Append accepted KV
- 5
5 Append accepted KV
- next6 On reject, resample once
- 6
6 On reject, resample once
- next7 Commit the corrected token
- 7
7 Commit the corrected token
Lesson map
Speculative Decoding
When decode is memory-bandwidth bound, speculative decoding spends spare FLOPs to cut sequential target steps. A draft proposes several tokens and the target verifies them in one forward. Acceptance rate decides whether you win or waste the step.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Draft model or extra heads"] b["2 Candidate tokens ahead"] c["3 Target forward in parallel"] d["4 Accept the matching prefix"] a -->|1 Draft model or extra heads| b b -->|2 Candidate tokens ahead| c c -->|3 Target forward in parallel| d
Sandbox: exact-match accept
The loop is greedy equality, not speculative sampling. Each drafted position matches with probability α until the first miss. The reported number is average tokens advanced per verify, with a floor of one.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The example [1, 2, 3, 4] against [1, 2, 9, 4] accepts 1, 2 and then the corrected 9. The trailing 4 is discarded even though it happens to match, because the chain already broke.
Pitfalls
Assume the interview formula (1 - α^(k+1)) / (1 - α) and ignore draft cost. At α = 0.8, compare k = 2 and k = 6. Then assume the draft itself costs a quarter of a target step. Does k = 6 still win? Write the inequality before you quote a vendor speedup.
Interview Q&A
Does speculative decoding change the output distribution?
Answer
Correct speculative sampling is designed to match the target distribution. Naive “accept if the greedy token is equal” only matches greedy decoding. If the product samples with temperature, say so, and say whether the reject path resamples from the target residual. A buggy rule biases outputs.
Why can speculation help when decode is memory-bound?
Answer
One target forward verifies multiple tokens and amortizes the KV read across every accepted output. You spend FLOPs you had spare while HBM was the limit. If the GPU is already compute-bound, that trade goes the wrong way.
How does it interact with continuous batching?
Answer
Sequences advance by a variable number of tokens per iteration. The scheduler and the KV pager must commit only accepted tokens and must size the iteration with that variable advance. See PagedAttention & Continuous Batching.
What does acceptance rate actually decide?
Answer
Return on the verify step. High α turns one target forward into several output tokens. Low α means you almost always reject immediately, so the draft was overhead. Measure α on your domain, not on the draft’s training set.
Medusa versus a separate draft model?
Answer
Medusa-style heads live on the target and propose future tokens without a second network. A separate draft is easier to swap and easier to mismatch. Both still need a verify step on the target. Training the heads is not this lesson.
What is tree attention doing, in one sentence?
Answer
It packs several draft branches into one target forward with a tree mask, so you verify more than a single chain. You pay a denser mask and extra transient state. It is still draft plus verify.
Why doesn’t speculation fix a long cold prompt?
Answer
The cold prompt is prefill, which is usually compute-bound. Speculation targets the decode tail. A prefix-cache hit is the lever for a repeated prompt. Speculation does not replace it.
What must happen to KV on a rejection?
Answer
Keep the accepted prefix and the single corrected token. Drop every drafted token after the first mismatch. If those tokens were written into the cache early, roll the blocks back. Otherwise later attention conditions on text the user never saw.
When is k too large?
Answer
When the marginal drafted token is rarely accepted, or when the draft’s own compute plus the fatter verify exceeds the tokens you gain. The formula grows with k only while α stays high. Past that, you are scoring tokens you will throw away.
Where should this sit in the mechanism ladder?
Answer
After paging and continuous batching, and only if the workload is decode-bound with a trustworthy draft. The hub classifier is LLM Inference Runtime — Prefill, KV, Batching & Parallelism. Do not turn it on to compensate for a bad batcher.
Go Deeper
- Fast Inference from Transformers via Speculative Decoding
- Medusa — multiple decoding heads
- EAGLE — feature-level draft
- vLLM speculative decoding
- Hub matrix: LLM Inference Runtime — Prefill, KV, Batching & Parallelism
- Next: Inference Parallelism — Tensor Parallel, Expert Parallel & Prefill/Decode Disaggregation