vLLM vs SGLang — Runtime Choice (comparative)
vLLM and SGLang share continuous batching and differ in the headline cache: paging versus a radix prefix tree. This page is a thin choice matrix. Prefix-tree internals, TensorRT engine builds, and playbook-length metric arithmetic stay in their own studies.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Which headline cache matches the traffic?
Prefer
Measure prefix reuse, then pick
High shared prefixes and structured flows point at SGLang. Heterogeneous multi-tenant chat that needs packing points at vLLM. An NVIDIA latency lock with an engine practice points at TensorRT-LLM.
- Both vLLM and SGLang do continuous batching.
- The split is paging-first versus radix-first, not batching versus no batching.
- A bakeoff on your traffic beats a logo.
Alternative
Pick the runtime you saw in a talk
SGLang with near-zero prefix reuse adds radix operations you never hit. vLLM plus a homegrown prefix cache in the app duplicates prefill. TensorRT engines churn if the weights are not stable.
- API compatibility is not feature parity.
- Tokens per second without a structured failure rate ships broken parsers.
- Training jobs are not this choice.
Overview
Lessons 2 through 5 are the shared mechanism stack: KV layout, paging, continuous batching, speculation, and serving parallelism. Product choice still matters, and the live site already has deep product pages. This lesson is comparative only, about one screen of product sketch plus a decision matrix.
Strict boundaries:
- No RadixAttention tree walk. SGLang — RadixAttention, Continuous Batching & Structured Generation.
- No engine build or quantization recipe. TensorRT-LLM — Engine Build, In-Flight Batching & Quantization.
- No playbook-length TTFT and cost redo. Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes.
- No JSON Schema tutorial. Structured Outputs / Constrained Decoding and the SGLang study.
One-screen sketches
vLLM (PagedAttention-centric). The origin story is PagedAttention plus continuous batching for high-throughput OpenAI-compatible serving. The model ecosystem is broad. Speculative decoding and multi-GPU tensor parallel show up as features. Check current docs before you memorize flags. Choose it for general serving when packing and ecosystem breadth dominate. Block-table depth is PagedAttention & Continuous Batching.
SGLang (RadixAttention-centric). The headline is a prefix tree for cross-request KV reuse, plus continuous batching and structured generation. It earns its keep when system prompts, multi-turn history, or routed templates repeat. Grammar-oriented serving is real. The tree and the API depth stay in the SGLang study.
TensorRT-LLM (latency-locked pointer). Ahead-of-time engine optimization, in-flight batching, and quantization. Choose it when the stack is NVIDIA, the latency budget is hard, and the team already runs engine builds. That lesson is the TensorRT-LLM study, not this page.
API surface
Both vLLM and SGLang expose familiar chat and completions HTTP for drop-in clients. Differences show up in extensions: structured or grammar endpoints, radix cache controls, and engine-specific flags. Do not treat “OpenAI-compatible” as identical behavior. Verify structured outputs and caching knobs on the version you will run.
Structured outputs, pointer only
If the requirement is JSON Schema or a grammar, use Structured Outputs / Constrained Decoding and SGLang — RadixAttention, Continuous Batching & Structured Generation. SGLang markets first-class programmatic flows. vLLM’s structured options move by release. Read the notes. Do not freeze an API from memory on this page.
Decision matrix
| You measured | Lean |
|---|---|
| High shared-prefix traffic (agents, RAG with a fixed system prompt) | SGLang, unless another constraint dominates |
| Heterogeneous multi-tenant chat, maximize packing | vLLM paging baseline |
| NVIDIA latency lock and an in-house quant recipe | TensorRT-LLM study |
| Multi-week training jobs | Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism, not these servers |
| ONNX, CPU, or edge export | ONNX & ONNX Runtime — Export, Graph Optimizers & Execution Providers |
| Unknown | Bake off prefix hit rate, TTFT p99, tokens per dollar, structured failure rate. Playbook metrics |
Flow
- 1
1 Measure prefix hit rate
- next2 Measure TTFT p99
- 2
2 Measure TTFT p99
- next3 Measure tokens per dollar
- 3
3 Measure tokens per dollar
- next4 Measure structured failures
- 4
4 Measure structured failures
- next5 High prefix leans SGLang
- 5
5 High prefix leans SGLang
- next6 NVIDIA lock leans TensorRT-LLM
- 6
6 NVIDIA lock leans TensorRT-LLM
- next7 Else general serving leans vLLM
- 7
7 Else general serving leans vLLM
Lesson map
vLLM vs SGLang — Runtime Choice (comparative)
vLLM and SGLang share continuous batching and differ in the headline cache: paging versus a radix prefix tree. This page is a thin choice matrix. Prefix-tree internals, TensorRT engine builds, and playbook-length metric arithmetic stay in their own studies.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Measure prefix hit rate"] b["2 Measure TTFT p99"] c["3 Measure tokens per dollar"] d["4 Measure structured failures"] a -->|1 Measure prefix hit rate| b b -->|2 Measure TTFT p99| c c -->|3 Measure tokens per dollar| d
The chain is a checklist, not a nested branch. Apply the first row that matches your constraint. A high prefix hit does not override a hard NVIDIA engine requirement if that requirement is real.
Sandbox: interview chooser
Order matches the toy: NVIDIA latency lock with TensorRT experience wins first, then prefix hit or structured need, otherwise vLLM. Thresholds are teaching numbers.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Failure modes
- Picking SGLang because of a conference talk, with near-zero prefix reuse. You pay radix complexity and never hit.
- Picking vLLM and then rebuilding a weak prefix cache in the application. You duplicate KV prefill the runtime should have shared.
- Forcing TensorRT engines before the weights are stable. Rebuild churn dominates any latency win. The engine lesson is the TensorRT-LLM study.
- Optimizing tokens per second and ignoring structured failure rate. Production parsers break.
- Assuming continuous batching alone holds TTFT. Admission control is still required. See PagedAttention & Continuous Batching.
| Runtime | Strength on this page | Where the depth lives |
|---|---|---|
| vLLM | Paging and continuous-batching maturity, broad models | This cluster’s paging lesson, plus vLLM docs |
| SGLang | Prefix reuse and structured flows | SGLang study |
| TensorRT-LLM | Engineered latency and quant | TensorRT-LLM study |
| This series | Bottleneck diagnosis | Lessons 2 through 5, not a substitute for a bakeoff |
Deep dive · What not to repeat from the playbook
Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes already scores TTFT against aggregate tokens per second, KV concurrency math, quantization cliffs, and managed APIs versus self-host. Use it when the interviewer wants the spreadsheet. Use this page when they want mechanism to product: paging versus radix versus an NVIDIA engine. Use the SGLang and TensorRT-LLM studies when they want implementation.
Pitfalls
Prefix hit rate 0.55, TTFT p99 acceptable, tokens per dollar tied, structured failure rate 8 percent on vLLM and 1 percent on SGLang. Which way do you lean, and which study do you open for the grammar? Now set prefix hit rate to 0.05 and structured failure to zero on both. What changes?
Interview Q&A
What is the core architectural contrast between vLLM and SGLang?
Answer
vLLM’s headline mechanism is PagedAttention for KV packing. SGLang’s headline mechanism is RadixAttention for cross-request prefix sharing. Both do continuous batching. Do not describe them as “batching versus no batching.”
When is TensorRT-LLM the better interview answer?
Answer
When the constraint is an NVIDIA deployment with ahead-of-time engine optimization and quantization, not when the question is about open paging or radix mechanics. Depth is TensorRT-LLM — Engine Build, In-Flight Batching & Quantization.
How do you avoid duplicating the Comparative Playbook?
Answer
Use the playbook for the metrics-led stack narrative. Use this page for mechanism-to-product mapping. Use the SGLang and TensorRT-LLM studies for implementation. If you start recomputing KV GiB at length, you are on the wrong page. That math is Prefill vs Decode & KV Cache Mechanics and the playbook.
High shared-prefix agent traffic. Which way do you lean?
Answer
SGLang, because the radix prefix cache is the lever, unless an NVIDIA latency lock or a missing model build overrides it. Confirm with prefix hit rate on your logs, not with a hope that prompts “look similar.”
Is an OpenAI-compatible route the same product?
Answer
No. Both servers speak chat and completions closely enough for many clients. Structured endpoints, cache controls, and error bodies differ. Test the extension you depend on.
Where do JSON Schema questions go?
Answer
Structured Outputs / Constrained Decoding, and the SGLang study for the runtime that markets programmatic flows. This page only notes that structured failure rate belongs in the bakeoff.
What is the failure mode of choosing SGLang with no prefix reuse?
Answer
You operate a radix cache that rarely hits. Complexity without the TTFT win. vLLM’s paging baseline is the simpler default for heterogeneous chat that does not share prefixes.
What is the failure mode of choosing vLLM and caching prompts in the app?
Answer
The application recomputes or stores prefixes the server could have reused, often incorrectly (unstable prefixes, no eviction story). You pay twice. If prefix reuse is the workload, pick a runtime whose headline cache is that reuse, then read the SGLang study.
A teammate asks which server to use for a multi-week FSDP run.
Answer
Neither. That is training. Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism. Serving tensor parallel from lesson 5 is still not FSDP.
Which four numbers do you collect before a bakeoff?
Answer
Prefix hit rate, TTFT p99, tokens per dollar, and structured failure rate. The playbook tells you how to read TTFT against aggregate throughput. This page tells you which product those numbers point at. Unknown numbers mean you do not get to pick a logo yet.
Go Deeper
- vLLM documentation
- SGLang documentation
- SGLang paper
- NVIDIA TensorRT-LLM documentation
- Hugging Face Text Generation Inference
- SGLang — RadixAttention, Continuous Batching & Structured Generation
- TensorRT-LLM — Engine Build, In-Flight Batching & Quantization
- Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes
- Structured Outputs / Constrained Decoding