SGLang — RadixAttention, Continuous Batching & Structured Generation
Most LLM serving pain is the runtime, not the model. Agents, chat, RAG, and multi-turn loops resend long shared prefixes. SGLang’s bet is prefix reuse via RadixAttention, plus continuous batching and a constrained-decoding API. This lesson is the runtime layer — it does not re-teach JSON Schema or CFGs.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Shared prefixes vs a generic server
Prefer
Radix prefix cache + continuous batch + in-runtime mask
Walk the tree, reuse the longest prefix, join the batch mid-flight, mask illegal tokens before sampling. Prefill work collapses when system prompts, RAG blocks, and prior turns overlap.
- A radix tree matches the longest shared prefix; a flat map only matches whole prompts.
- Continuous batching keeps the GPU busy; prefix reuse is a separate win.
- Constraints compose with cache and batch because the runtime owns the mask.
Alternative
Recompute the prefix, then regex-validate JSON
You pay prefill N times for one conversation, then retry invalid JSON. The “fix” looks like more GPUs.
- Zero overlap traffic gets little from the radix tree.
- Post-hoc validation plus retries wastes tokens and adds tail latency.
- vLLM-style paging still wins on coverage when prefixes do not repeat.
One SGLang iteration
Lookup, reuse, join, mask, sample, cache. Schema/CFG depth stays in Structured Outputs.
- 1
Ingest and tokenize
Client request plus the shared prefix. Tokenize once. - 2
Radix lookup
Match the longest existing token prefix; reuse those KV blocks. - 3
Join the batch
Continuous batching admits the request mid-flight. - 4
Mask, sample, emit
Apply the constraint as a logit mask, sample, stream, insert new radix nodes.
Overview
Most LLM serving pain is not the model — it is the runtime. Agents, chat, RAG, and multi-turn tool loops resend long, nearly identical prefixes. A naive server recomputes that prefix on every request.
SGLang’s bet: prefix reuse is the dominant optimization for LLM applications, exposed through RadixAttention, continuous batching, and a constrained-decoding API. This lesson does not re-teach JSON Schema or context-free grammars — see Structured Outputs.
SGLang runtime for LLM apps
SGLang is a serving runtime plus a front-end language for generation programs. Instead of one prompt → one completion, you write a program with control flow, forks, joins, and parallel branches. The runtime schedules those branches against a shared KV cache. The runtime is the product; the DSL is the ergonomic surface.
RadixAttention and the prefix cache
RadixAttention keeps KV cache blocks in a radix tree keyed by token sequences. A new request walks the tree, matches the longest existing prefix, and reuses those blocks instead of recomputing them. Divergent suffixes branch off; eviction is LRU-ish over tree nodes.
Consequence: with high prefix overlap (chat history, shared system prompts, multi-sample fan-out), prefill work collapses. With zero overlap, it is close to a plain cache and buys little.
Continuous batching
Requests are scheduled per-iteration, not per-request: finished sequences leave the batch and queued ones join mid-flight. Combined with paged KV memory, this keeps the GPU busy under bursty, heterogeneous traffic.
Continuous batching is now table stakes. SGLang’s differentiator is that it batches while sharing radix prefixes across the batch.
Constrained / structured decoding hooks
SGLang accepts a grammar / regex / JSON-schema constraint and applies it as a logit mask at each step, so the sampler can only emit tokens that keep the output valid. Same family of technique as Structured Outputs — here the point is only that the runtime owns the mask, so it composes with prefix caching and batching rather than being a post-hoc validator.
Streaming the constrained tokens: Streaming Structured Output. Multi-turn tool loops are a traffic pattern, not a tools lesson — tool calling vs structured outputs if you need that fork.
Vendor-reported: constrained decoding can add per-step overhead and, for strict grammars, reduce throughput.
OpenAI-compatible serving
SGLang ships an OpenAI-compatible HTTP surface (/v1/chat/completions, /v1/completions, streaming SSE). Existing clients point at it by changing base_url. Extra SGLang-specific fields (constraints, n, cache hints) ride along as extensions. OpenAI-compatible ≠ OpenAI-identical for extensions and error shapes.
Light comparison to vLLM / PagedAttention
vLLM popularized PagedAttention — KV cache as fixed-size pages with a block table, which kills fragmentation. SGLang reuses paged KV but adds the radix tree for cross-request prefix sharing and a program-level scheduler.
Rule of thumb: heavy shared prefixes and branching programs → SGLang; broad model coverage, mature ecosystem, and simple high-throughput single-turn serving → vLLM. Both are vendor-reported to be fast; benchmark on your traffic mix. This is a comparator, not a vLLM lesson.
Wrong-stack symptoms
- Every request recomputes a 4k-token system prompt → prefix cache disabled or keys not stable.
- GPU utilization sawtooths with idle gaps → static batching, not continuous.
- You regex-validate JSON after generation and retry on failure → constraints belong in the sampler.
- You wrote a bespoke async fan-out layer above a server that already schedules branches.
Comparative: SGLang vs a vLLM-style baseline
| Dimension | SGLang | vLLM-style baseline |
|---|---|---|
| Shared-prefix workloads | Strong (radix reuse) | Weaker without explicit prefix caching |
| Program-level branching | Native | Usually client-side |
| Constrained decoding | In-runtime logit mask | Varies by version / plugin |
| Model / quant coverage | Narrower | Broader |
| Ops familiarity | Newer surface | Widely deployed |
Wrong stack, concretely: picking SGLang for a single-turn, zero-overlap workload wastes its advantage; picking a plain server for a 20-turn agent loop multiplies prefill cost. The fix looks like “buy more GPUs” when it is really “stop recomputing the prefix.”
Flow
- 1
1 Client request + prefix
- next2 Tokenize
- 2
2 Tokenize
- next3 Radix tree lookup
- 3
3 Radix tree lookup
- next4 Reuse matched KV blocks
- 4
4 Reuse matched KV blocks
- next5 Continuous batch join
- 5
5 Continuous batch join
- next6 Logit mask from constraint
- 6
6 Logit mask from constraint
- next7 Sample token
- 7
7 Sample token
- next8 Stop or decode again
- 8
8 Stop or decode again
- next9 Stream response
- 9
9 Stream response
- next10 Insert radix nodes
- 10
10 Insert radix nodes
Lesson map
SGLang — RadixAttention, Continuous Batching & Structured Generation
Most LLM serving pain is the runtime, not the model. Agents, chat, RAG, and multi-turn loops resend long shared prefixes. SGLang’s bet is prefix reuse via RadixAttention, plus continuous batching and a constrained-decoding API. This lesson is the runtime layer — it does not re-teach JSON Schema or CFGs.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Client request + prefix"] b["2 Tokenize"] c["3 Radix tree lookup"] d["4 Reuse matched KV blocks"] a -->|1 Client request + prefix| b b -->|2 Tokenize to 3 Radix tree lookup| c c -->|3 Radix tree lookup| d
Sandbox: mock radix cache (Python)
No GPU, no real KV. Shared prefixes increment reused_tokens.
Press Run. Snippets must be self-contained — no network, files, or native modules.
TypeScript client shape
Production POSTs to /v1/chat/completions. The playground mocks the transport.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Deep dive · Structured Outputs overlap — pointer only
The mask lives in the sampler. How schemas compile to illegal-token sets, streaming partial JSON, and tools vs JSON is not this lesson: Structured Outputs, Streaming Structured Output, tool calling vs structured outputs.
Pitfalls
Take a 4k system prompt. Put a request_id on line 1. What happens to radix hits? Where would you move the id so the shared prefix stays stable?
Interview Q&A
Why a radix tree instead of a flat prefix cache?
Answer
A flat map only matches whole prompts. A radix tree matches the longest shared prefix across many requests and lets divergent suffixes share ancestors — the chat / agent pattern.
Does continuous batching alone give you prefix reuse?
Answer
No. Continuous batching improves GPU occupancy; prefix reuse is a separate memory-and-compute optimization. SGLang does both, and they compose.
Where should a JSON constraint live?
Answer
In the sampler as a logit mask, so invalid tokens are never emitted. Post-hoc validation plus retries wastes tokens and adds tail latency. Schema/CFG depth: Structured Outputs.
When would you choose vLLM over SGLang?
Answer
When model coverage, quantization breadth, or ecosystem maturity dominates, or when traffic has near-zero prefix overlap and you just need steady throughput. Comparator only — do not recap PagedAttention internals here.
Go Deeper
- SGLang documentation — runtime, RadixAttention, constrained decoding
- SGLang paper
- vLLM docs — PagedAttention and continuous batching, baseline comparison
- Next: TensorRT-LLM — Engine Build, In-Flight Batching & Quantization