LLM Inference Runtime — Prefill, KV, Batching & Parallelism
Senior interviews fail when they name a serving product before the bottleneck. This hub is the mechanism ladder under the live LLM Inference Engines cluster: prefill versus decode, KV bytes, PagedAttention, continuous batching, speculative decode, then serving tensor parallel, expert parallel, and prefill/decode disaggregation.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
What are you diagnosing?
Prefer
Name the binding mechanism, then the product
Prefill, decode bandwidth, prefix reuse, speculation, and serving collectives are different knobs. The product runtime is the last box on the board.
- Long cold prompts point at prefill capacity, chunked prefill, or a prefill/decode split.
- Many short generations point at KV packing and an iteration-level scheduler.
- Shared system prompts point at prefix reuse. That tree lives in the SGLang study.
Alternative
Add tensor parallel because the dashboard looks busy
An 8-way shard on a tiny chat model pays an all-reduce on every decode step. TTFT gets worse and the real issue was fragmentation or admission control.
- Vendor charts are not your prompt mix.
- Training topology is not the serving topology.
- Continuous batching without admission control blows the TTFT SLO.
Interview order of operations
Mechanism first. Product second. Measure on your traffic.
- 1
Write the shape
Prompt length, generation length, concurrency, prefix hit rate, draft accept rate, dense versus MoE. - 2
Name the bound
Compute-bound prefill, memory-bound decode, collective floor, or expert hotspot. - 3
Pick the lever
Paging, continuous batching, prefix cache, speculation, serving TP, EP, or a prefill/decode split. - 4
Point at the product study
vLLM, SGLang, or TensorRT-LLM. Metrics arithmetic stays in the Comparative Playbook.
Overview
The live AI/ML topic already teaches product runtimes: SGLang — RadixAttention, Continuous Batching & Structured Generation, TensorRT-LLM — Engine Build, In-Flight Batching & Quantization, the engines hub LLM Inference Engines & Distributed Training — SGLang, TensorRT-LLM, ONNX & PyTorch Parallelism, Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes, Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism, and ONNX & ONNX Runtime — Export, Graph Optimizers & Execution Providers.
Those pages assume a mechanism stack. This cluster teaches that stack:
- One forward per new token after the prompt.
- Prefill is usually compute-bound. Decode is usually memory-bandwidth-bound because it streams KV.
- PagedAttention stores KV in blocks with a table, including copy-on-write for beams.
- Continuous batching changes membership every iteration.
- Speculative decode drafts tokens and verifies them when the draft is trustworthy.
- Serving tensor parallel, expert parallel, and prefill/decode disaggregation scale inference. They are not a training 3D grid.
Ladder map
| Step | Mechanism | What it buys |
|---|---|---|
| Autoregressive core | Prefill, then one token at a time | Correct generation loop |
| KV cache | Store K and V per layer | Decode does not recompute past attention |
| PagedAttention | Block table, non-contiguous KV, copy-on-write | Packing and beam forks |
| Continuous batching | Iteration-level schedule | Occupancy across mixed lengths |
| Prefix / radix cache | Share prompt prefixes | TTFT on repeated system prompts |
| Speculative decode | Draft plus verify | Lower time per output token when acceptance is high |
| Serving tensor parallel | Shard weights and KV heads | Fit and single-stream latency on a fast fabric |
| Expert parallel | Route tokens, all-to-all | Sparse MoE capacity |
| Prefill/decode split | Separate pools | Independent TTFT and tokens-per-second knobs |
Prefix trees are production depth in the SGLang study. This hub only names the lever.
When each bottleneck dominates
- Long prompts, cold cache. Time-to-first-token and prefill FLOPs dominate. Add prefill capacity, chunk the prefill, or split pools.
- Many concurrent short generations. KV memory and scheduler fragmentation dominate. PagedAttention plus continuous batching.
- Shared system prompts or multi-turn chat. Prefix hit rate dominates. Radix / prefix cache. See the SGLang study.
- Decode-bound, draft model trustworthy. Speculative decoding.
- Single-request latency on several GPUs. Serving tensor parallel. Not a training mesh.
- Sparse mixture-of-experts. Expert parallel plus load balance.
- TTFT SLO fights tokens-per-second on the same GPUs. Prefill/decode disaggregation.
- Hard latency lock on NVIDIA. TensorRT-LLM — Engine Build, In-Flight Batching & Quantization.
- Structured JSON or grammar at scale. Structured Outputs / Constrained Decoding and the SGLang study. Pointer only.
Deliberately not re-taught
| Already taught | Where |
|---|---|
| RadixAttention insert and match | SGLang — RadixAttention, Continuous Batching & Structured Generation |
| Engine build, in-flight batching productization, quantization recipes | TensorRT-LLM — Engine Build, In-Flight Batching & Quantization |
| Playbook-length TTFT, tokens-per-second, and cost arithmetic | Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes |
| DDP, FSDP, training tensor and pipeline grids | Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism |
| ONNX export and execution providers | ONNX & ONNX Runtime — Export, Graph Optimizers & Execution Providers |
| JSON Schema constrained decoding | Structured Outputs / Constrained Decoding |
Runtime ladder
Teaching order. A deployment may stop at paging.
Flow
- 1
1 Prefill the prompt
- next2 Decode one token
- 2
2 Decode one token
- next3 KV cache holds K and V
- 3
3 KV cache holds K and V
- next4 PagedAttention blocks
- 4
4 PagedAttention blocks
- next5 Continuous batching
- 5
5 Continuous batching
- next6 Prefix cache when shared
- 6
6 Prefix cache when shared
- next7 Speculative draft plus verify
- 7
7 Speculative draft plus verify
- next8 Serving tensor parallel
- 8
8 Serving tensor parallel
- next9 Expert parallel for MoE
- 9
9 Expert parallel for MoE
- next10 Prefill and decode pools
- 10
10 Prefill and decode pools
Lesson map
LLM Inference Runtime — Prefill, KV, Batching & Parallelism
Senior interviews fail when they name a serving product before the bottleneck. This hub is the mechanism ladder under the live LLM Inference Engines cluster: prefill versus decode, KV bytes, PagedAttention, continuous batching, speculative decode, then serving tensor parallel, expert parallel, and prefill/decode disaggregation.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Prefill the prompt"] b["2 Decode one token"] c["3 KV cache holds K and V"] d["4 PagedAttention blocks"] a -->|1 Prefill the prompt| b b -->|2 Decode one token| c c -->|3 KV cache holds K and V| d
Wrong bottleneck
Eight-way tensor parallel on chatty short prompts spends the step on the collective. Paging, packing, and a prefix hit raise tokens per second without moving the TTFT SLO the wrong way.
Flow
- 1
1 Misread short chat as TP
- next2 Shard a tiny model 8-way
- 2
2 Shard a tiny model 8-way
- next3 All-reduce tax beats compute
- 3
3 All-reduce tax beats compute
- next4 TTFT gets worse
- 4
4 TTFT gets worse
- next5 Page the KV cache instead
- 5
5 Page the KV cache instead
- next6 Pack decode-bound requests
- 6
6 Pack decode-bound requests
- next7 Hit the shared system prompt
- 7
7 Hit the shared system prompt
- next8 Higher TPS at the same TTFT
- 8
8 Higher TPS at the same TTFT
Sandbox: bottleneck classifier
No GPU and no network. First matching rule wins. It is an interview framing, not a profiler. A decode-bound request with a high prefix hit is classified as decode before prefix, because the SLO breach is checked first.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Mechanism levers
| Lever | Gain | Cost |
|---|---|---|
| PagedAttention | High memory use, copy-on-write beams | Bookkeeping and block-size tuning |
| Continuous batching | Occupancy across mixed lengths | Fairness and starvation risk |
| Prefix / radix cache | Large TTFT wins on shared prompts | Eviction policy. See the SGLang study |
| Speculative decode | Lower time per output token when acceptance is high | Wasted FLOPs when the draft is poor |
| Serving tensor parallel | Fit large models, cut latency | Collective floor on small batches |
| Expert parallel | Sparse capacity | All-to-all and expert imbalance |
| Prefill/decode split | Independent TTFT and throughput scaling | KV transfer path and ops complexity |
Deep dive · Product cluster, one sentence each
SGLang — RadixAttention, Continuous Batching & Structured Generation owns the prefix tree. TensorRT-LLM — Engine Build, In-Flight Batching & Quantization owns engine build and quantization. Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes owns stack metrics. Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism owns the training grid. ONNX & ONNX Runtime — Export, Graph Optimizers & Execution Providers owns portable export. This series does not repeat those lessons.
Pitfalls
Ticket A: 8k-token prompts, 32-token answers, p99 time-to-first-token over a second, eight concurrent users. Ticket B: 200-token prompts, 500-token answers, 64-way concurrency, slow time per output token. Ticket C: the same system prompt on 70% of calls. Which lever do you reach for first, and which existing study do you open second?
Interview Q&A
Why is decode often memory-bandwidth bound while prefill is compute bound?
Answer
Prefill attends over the full prompt with large matrix multiplies, so arithmetic intensity is high. Decode does a thin multiply for one new token but must stream the entire KV cache from high-bandwidth memory. Bytes moved dominate FLOPs. Extra tensor-core throughput helps prefill more than it helps decode.
How does PagedAttention relate to continuous batching?
Answer
Continuous batching admits and finishes requests every iteration instead of waiting for a static batch. Paging lets the scheduler allocate and free KV in fixed blocks without reserving a contiguous max-length slab. They compose. Depth is the next two lessons.
When would you pick SGLang over vLLM at the mechanism level?
Answer
High prefix reuse and structured generation point at a RadixAttention-centric runtime. Read SGLang — RadixAttention, Continuous Batching & Structured Generation. General OpenAI-compatible serving with a mature paging ecosystem points at vLLM. A latency-locked NVIDIA engine points at TensorRT-LLM — Engine Build, In-Flight Batching & Quantization. The thin matrix is lesson 6.
Is inference tensor parallel the same as training tensor parallel?
Answer
The weight-shard idea is the same. The goal is not. Serving tensor parallel optimizes latency and throughput, splits KV heads, and runs small collectives on the decode critical path. Training tensor parallel sits in a 3D plan with FSDP and pipeline stages. Open Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism for that grid. Do not re-derive it here.
What is the default mechanism baseline?
Answer
Paged KV plus an iteration-level scheduler. Prefix caching and speculative decoding are amplifiers you turn on when hit rate or acceptance rate earns them. Serving parallelism comes after you know the single-pool bottleneck.
A team added 8-way tensor parallel to a 7B chat service and TTFT got worse. What happened?
Answer
The all-reduce tax exceeded the compute saved. Short prompts on a slow fabric are the classic miss. Measure paging, admission control, and prefix hit rate before adding ranks.
Where do playbook metrics live?
Answer
Time-to-first-token, tokens per second, KV capacity, and dollars per million tokens are defined just enough here to name a bottleneck. The scored stack comparison is Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes.
Why mention ONNX and structured outputs on a mechanism hub?
Answer
So you know which question to refuse. Portable CPU or edge export is ONNX & ONNX Runtime — Export, Graph Optimizers & Execution Providers. Grammar and JSON Schema constraints are Structured Outputs / Constrained Decoding. Neither is a KV-layout lesson.
When does prefill/decode disaggregation show up?
Answer
When one GPU pool cannot hit a strict time-to-first-token SLO and a high decode concurrency SLO at once. Long prefills stall decodes. Split pools, then pay for KV transfer. Lesson 5.
What should you say if the interviewer asks you to draw a RadixAttention tree?
Answer
Name the lever (cross-request prefix reuse) and open the SGLang study. Insert, match, and eviction policy are that page, not this one.
Go Deeper
- vLLM documentation — PagedAttention and continuous batching
- PagedAttention paper
- SGLang documentation — RadixAttention overview
- Hugging Face Text Generation Inference
- NVIDIA TensorRT-LLM — in-flight batching and disaggregated serving live in the engine docs, not this hub
- Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes
- Next: Prefill vs Decode & KV Cache Mechanics