Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes
The serving stack is where model quality meets a bill and a p99. Interviews test TTFT vs TPS, KV capacity math, quantization cliffs, and a defensible matrix — not “do you know vLLM.” This playbook ties the cluster together.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Which constraint binds?
Prefer
Pick the binding constraint, then the stack
Iteration speed, latency SLO, cost, or hardware portability. TRT-LLM wins raw NVIDIA perf; it loses when weights churn weekly. SGLang wins prefix-heavy agents. ONNX wins mixed silicon. Managed APIs win when utilization is low.
- Throughput knobs (continuous batching, paged KV) buy aggregate TPS.
- Latency knobs (chunked prefill, prefix cache) buy TTFT. You rarely get both from one knob.
- Self-host vs API is arithmetic: GPU $/hour ÷ achieved TPS vs blended $/1M.
Alternative
Ship the stack the team already knows
TRT-LLM for a weekly retrain. Raw PyTorch at 200 QPS. ONNX Runtime on 8×H100 because someone exported once. Familiarity is not a latency SLO.
- Vendor-reported speedups move with hardware, batch, and architecture.
- Sizing KV for max context × max concurrency is how you OOM.
- INT4 can hold general evals and still break JSON / code / math.
Decision order interviewers like
Churn first, then SLO, then hardware and ops depth. Measure on your mix.
- 1
Is the model iterating weekly?
Yes → raw PyTorch or a managed API. Engine CI will not keep up. - 2
Is the latency SLO tight?
Loose → SGLang / vLLM-class serving. Tight + NVIDIA → TRT-LLM or SGLang. - 3
What silicon?
Non-NVIDIA or mixed → ONNX Runtime. NVIDIA + GPU team + steady load → self-host and pin versions. - 4
Do the arithmetic
GPU day rate vs API $/1M at your real prompt / completion mix. Pin the version matrix.
Overview
The serving stack is where model quality meets a bill and a p99. Two teams with identical weights can differ 5–10× on cost per million tokens purely from runtime choice, scheduler behavior, and cache policy.
Interviews for AI/ML systems roles test this directly: not “do you know vLLM” but “given a 3-second TTFT SLA, 8×H100, and an MoE checkpoint, what do you ship and why?” This playbook gives you the metrics vocabulary, the failure taxonomy, and a defensible decision matrix.
TTFT, TPS, cost, GPU memory
- TTFT (time to first token): queue delay + prefill compute. Dominated by prompt length, batch composition, and whether chunked prefill is enabled. Report p50 and p99; p99 TTFT is usually a scheduling problem, not a kernel problem.
- TPS (tokens per second): separate per-request decode TPS (latency) from aggregate / system TPS (throughput). They trade off directly — larger batches raise aggregate TPS and hurt per-request TPS.
- Cost: normalize to $/1M output tokens and $/1M input tokens at a fixed SLA. Include GPU hourly rate ÷ achieved aggregate TPS, plus idle capacity held for burst headroom.
- GPU memory: weights + KV cache + activations + fragmentation. KV cache is the elastic term: bytes ≈
2 × layers × kv_heads × head_dim × dtype_bytes × tokens. Concurrency ceiling ≈(free_mem − safety_margin) / per_token_kv_bytes.
Rule of thumb: throughput optimizations (continuous batching, paged KV, speculative decode) buy aggregate TPS; latency optimizations (chunked prefill, prefill/decode disaggregation, prefix caching) buy TTFT. You rarely get both from one knob.
Failure modes
KV cache OOM. Not a leak — a capacity planning miss. Symptoms: OOM under load spikes, preemption thrash, requests queued behind cache eviction. Fixes: paged attention, KV quantization (FP8), radix / prefix cache sharing, admission control, and a hard max-concurrency cap derived from the formula above. Never size KV for peak concurrency at max context; size for expected concurrency at weighted-average context.
Batching starvation. The scheduler waits for a full batch that never arrives, or a long prefill blocks decodes. Symptoms: low GPU utilization at low QPS, TTFT spikes correlated with long prompts. Fixes: continuous batching with token budgets, chunked prefill, separate prefill / decode pools.
Quantization quality cliffs. INT8 / FP8 usually survive; INT4 / AWQ / GPTQ degrade unevenly — often fine on general chat, sharp on code, math, JSON-schema adherence, and long-context retrieval. Always gate on task-specific evals, not perplexity alone. Cliffs appear as format violations before they appear as wrong answers. Format validity light-link: Structured Outputs and tool calling vs structured outputs — do not recap schemas here.
Ops engine rebuilds / CUDA versions. Prebuilt wheels lag new GPU arch and CUDA / PyTorch combos. Symptoms: silent fallback to slower kernels, or a 90-minute JIT / AOT build in CI. Pin a tested matrix (driver, CUDA, PyTorch, engine) and treat it as a versioned artifact. Rebuild cost is a real tax on iteration speed — see TensorRT-LLM.
Interview decision matrix
| Stack | Best when | Watch out |
|---|---|---|
| SGLang | High-concurrency serving, prefix-heavy workloads (RAG, agents), RadixAttention reuse | Newer ecosystem; validate custom model support early |
| TensorRT-LLM | Fixed model, max perf on NVIDIA, willing to compile engines | Build / pin complexity; engine tied to GPU arch + TRT version; speedups are vendor-reported |
| ONNX Runtime | Cross-platform / edge, CPU or mixed hardware, non-NVIDIA targets | LLM decode perf generally behind GPU-native engines on NVIDIA |
| Raw PyTorch | Research, custom kernels, unusual architectures, fastest iteration | You build batching, KV cache, paging yourself — only sane at low QPS |
| Managed API | Spiky traffic, no GPU team, fastest time-to-prod, frontier models | Per-token cost at scale, data residency, rate limits, less control |
Managed APIs beat self-host when sustained utilization is low (roughly under 30%), the model is commodity, or the team lacks on-call GPU depth. Self-host wins at high steady utilization, custom / fine-tuned weights, strict latency SLOs, or data-locality requirements. The crossover is arithmetic, not ideology.
Depth pointers: SGLang, TensorRT-LLM, ONNX Runtime, Distributed PyTorch for the train-time grid.
Pros / cons and the wrong-stack smell
Wrong-stack signals: choosing TRT-LLM for a model you retrain weekly; choosing raw PyTorch for 200 QPS; choosing a managed API for a 40B fine-tune with nightly eval loops; choosing ONNX Runtime on 8×H100 because the team already knew it.
The right question is not “which is fastest” but which constraint binds — iteration speed, latency SLO, cost, or hardware portability.
Flow
- 1
1 Model fixed or iterating?
- next2 Weekly retrain: PyTorch or API
- 2
2 Weekly retrain: PyTorch or API
- next3 Tight latency SLO?
- 3
3 Tight latency SLO?
- next4 NVIDIA + team: TRT-LLM
- 4
4 NVIDIA + team: TRT-LLM
- next5 Prefix-heavy: SGLang
- 5
5 Prefix-heavy: SGLang
- next6 Mixed silicon: ONNX Runtime
- 6
6 Mixed silicon: ONNX Runtime
- next7 No GPU team: managed API
- 7
7 No GPU team: managed API
- next8 Pin versions and measure
- 8
8 Pin versions and measure
Lesson map
Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes
The serving stack is where model quality meets a bill and a p99. Interviews test TTFT vs TPS, KV capacity math, quantization cliffs, and a defensible matrix — not “do you know vLLM.” This playbook ties the cluster together.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Model fixed or iterating?"] b["2 Weekly retrain: PyTorch or API"] c["3 Tight latency SLO?"] d["4 NVIDIA + team: TRT-LLM"] a -->|1 Model fixed or iterating?| b b -->|2 Weekly retrain: PyTorch or API| c c -->|3 Tight latency SLO?| d
Linear map of the interview order. vLLM-class serving sits with SGLang on the “not weekly-retrain, SLO not ultra-tight” rung — comparator, not a full lesson.
Sandbox: stack scorer (Python)
Pure Python. Swap the weights and the ranking flips — that is the point.
Press Run. Snippets must be self-contained — no network, files, or native modules.
TRT-LLM wins on pure latency / throughput; it loses hard on ops and flexibility once retraining cadence enters the picture.
TypeScript: cost per 1M tokens
Press Run. Snippets must be self-contained — no network, files, or native modules.
Deep dive · Sibling depth — pointer only
Radix prefix reuse: SGLang. Engine rebuilds: TensorRT-LLM. Portable graphs: ONNX Runtime. Train-time grids: Distributed PyTorch. Format cliffs: Structured Outputs — do not recap JSON Schema. vLLM is a comparator for paged KV and continuous batching.
Pitfalls
What do you inspect first — kernels or the scheduler? Name three knobs, then compute a KV concurrency cap for 80 GB free, 0.5 MB per token, 20% safety margin.
Interview Q&A
TTFT p99 is 6s, p50 is 400ms. What do you look at first?
Answer
Batching and queueing, not kernels. Check max batch size, whether long prefills block decodes, chunked prefill settings, and admission control. A huge p50 / p99 gap is a scheduler signature.
KV cache OOM at peak. Walk me through triage.
Answer
Compute per-token KV bytes, derive max concurrency at weighted-average context, compare to configured limits. Then enable paged attention and prefix caching, quantize KV to FP8, cap max context, and add admission control. Treat it as capacity math, not a bug hunt.
When is a managed API the right answer over self-hosting?
Answer
Low sustained utilization, no GPU on-call, commodity model, spiky traffic, or a need to ship this week. The arithmetic: compare GPU $/hour ÷ achieved aggregate TPS against the API’s blended $/1M tokens at your real prompt / completion mix.
You quantized to INT4 and general evals held. What else do you test?
Answer
Structured output validity (JSON schema, tool calls), long-context retrieval, code and math, and refusal / format-compliance rates. Quality cliffs show up in strict-format tasks first. Keep an FP8 or FP16 fallback behind a flag. Schema depth: Structured Outputs — pointer only.
Go Deeper
Public official docs only:
- SGLang docs
- NVIDIA TensorRT-LLM docs
- ONNX Runtime docs
- PyTorch distributed docs
- vLLM docs — continuous batching and PagedAttention, comparator only
- Return to the hub