LLM Inference Engines & Distributed Training — SGLang, TensorRT-LLM, ONNX & PyTorch Parallelism
Senior SWE/ML interviews ask you to pick a serving stack and a parallelism strategy under latency SLOs, cost per token, hardware mix, and team maturity. Mis-picking SGLang vs TensorRT-LLM vs ONNX Runtime vs raw PyTorch — or confusing DDP with FSDP — produces symptoms that look like “the model is slow” but are architectural. This hub maps the cluster.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
What are you optimizing?
Prefer
Name SLO, hardware, and churn — then pick the engine
Serving wants latency, concurrency, and cost per token. Training wants throughput and fault tolerance. A model can use different parallelism at train time and serve time.
- SGLang when structured generation, prefix reuse, or rapid iteration dominate.
- TensorRT-LLM when NVIDIA-only max perf is worth engine CI.
- ONNX Runtime for portable CPU/GPU/edge; raw PyTorch for research and custom ops.
Alternative
Copy the vendor chart, then debug “the model is slow”
Low GPU util, long-context OOM, and concurrency cliffs are usually batching or KV — not the weights. Vendor-reported speedups are not your p99.
- Benchmarking batch size 1 is not production.
- Training topology is not automatically the serving topology.
- Managed APIs hide the engine when time-to-market beats unit cost.
Interview order of operations
Requirements first. Engine second. Parallelism third. Measure on your traffic.
- 1
Write the SLO
Latency, concurrency, cost per token, hardware mix, model churn. - 2
Pick the engine
SGLang, TensorRT-LLM, ONNX Runtime, or raw PyTorch — see the matrix below. - 3
Pick parallelism
Serving: TP/PP and KV distribution. Training: DDP/FSDP, then TP/PP. - 4
Measure, do not trust the chart
Vendor-reported TTFT and tokens/sec need a replay of your prompt mix.
Overview
Senior SWE/ML systems interviews ask you to choose a serving stack and a parallelism strategy under real constraints. This hub builds the decision reflexes; siblings own the depth.
You should be able to:
- Draw training vs serving without mixing gradient sync into a KV-cache story.
- Defend one engine from the matrix, including when it is wrong.
- Point at SGLang, TensorRT-LLM, ONNX Runtime, DDP/FSDP, and the playbook as separate full lessons.
Training vs serving
Training optimizes for throughput over large batches, high memory, and fault tolerance. Serving optimizes for latency, concurrency, and cost per token.
Training parallelism (DDP / FSDP / TP / PP) splits gradients, parameters, or layers. Serving engines split the KV cache and batch dynamically. Do not assume the training topology is the serving topology.
Why specialized engines exist
- KV cache — avoids recomputing past keys/values. Paged KV (vLLM / PagedAttention as a comparator) reduces fragmentation.
- Continuous batching — admits new requests mid-flight; raises GPU utilization.
- Kernel fusion — merges ops to cut memory traffic and launch overhead.
- Quantization — FP8 / INT8 / INT4 reduces memory and can raise throughput, sometimes with accuracy trade-offs.
SGLang and TensorRT-LLM publish aggressive latency/throughput numbers; treat those as vendor-reported until reproduced on your workload.
Decision matrix
| Stack | Best for | Avoid when |
|---|---|---|
| SGLang | Structured generation, RadixAttention prefix reuse, rapid prototyping | You need NVIDIA-only max perf and have TensorRT expertise |
| TensorRT-LLM | Max NVIDIA throughput / latency, production engines | Non-NVIDIA hardware, fast-changing models |
| ONNX Runtime | Cross-platform, edge, CPU/GPU variety, graph optimization | You need bleeding-edge LLM kernels |
| Raw PyTorch | Research, custom ops, small scale | You need high-concurrency serving |
Where DDP / FSDP / TP / PP fit
- DDP — replicate the model, split the batch; simple, bandwidth-heavy.
- FSDP — shard parameters, gradients, optimizer states; large models on many GPUs.
- TP — split tensors within layers; high bandwidth, low latency.
- PP — split layers into stages; pipeline bubbles, good for very deep models.
Serving engines focus on inference TP/PP and KV distribution, not training. Depth: Distributed PyTorch.
Wrong-stack symptoms
- Low GPU utilization with high latency → batching or KV cache.
- OOM at long context → KV growth or missing quantization.
- Great single-stream, terrible concurrency → no continuous batching.
- Cross-platform portability pain → NVIDIA-specific engines.
Managed APIs
Managed endpoints (OpenAI, Anthropic, Bedrock) hide engine choice. Use them when time-to-market beats unit cost; self-host when you need control, privacy, or cost at scale. Arithmetic lives in the playbook.
Comparative pros / cons
SGLang: strong structured output and prefix caching; smaller ecosystem than TensorRT-LLM.
TensorRT-LLM: top NVIDIA perf; steep build and engine management.
ONNX Runtime: portable; fewer LLM-specific optimizations.
Raw PyTorch: flexible; you build batching and KV yourself.
Wrong choice: you either over-engineer or bottleneck at the engine layer, and retries/rollbacks cost weeks.
Flow
- 1
1 Name latency SLO
- next2 Name hardware mix
- 2
2 Name hardware mix
- next3 Name model churn
- 3
3 Name model churn
- next4 Pick serving engine
- 4
4 Pick serving engine
- next5 SGLang: structure + prefix
- 5
5 SGLang: structure + prefix
- next6 TRT-LLM: NVIDIA max perf
- 6
6 TRT-LLM: NVIDIA max perf
- next7 ONNX Runtime: portable
- 7
7 ONNX Runtime: portable
- next8 PyTorch: research / train
- 8
8 PyTorch: research / train
- next9 Add TP/PP or DDP/FSDP
- 9
9 Add TP/PP or DDP/FSDP
- next10 Deploy and measure
- 10
10 Deploy and measure
Lesson map
LLM Inference Engines & Distributed Training — SGLang, TensorRT-LLM, ONNX &
Senior SWE/ML interviews ask you to pick a serving stack and a parallelism strategy under latency SLOs, cost per token, hardware mix, and team maturity. Mis-picking SGLang vs TensorRT-LLM vs ONNX Runtime vs raw PyTorch — or confusing DDP with FSDP — produces symptoms that look like “the model is slow” but are architectural. This hub maps the cluster.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Name latency SLO"] b["2 Name hardware mix"] c["3 Name model churn"] d["4 Pick serving engine"] a -->|1 Name latency SLO| b b -->|2 Name hardware mix| c c -->|3 Name model churn| d
Linear map, not a nested diamond tree. The playbook owns the scored matrix.
Sandbox: choose an engine (Python)
No GPU, no API keys. The function encodes the matrix above.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same request shape (TypeScript)
jsonSchema is a constrained-decoding hook, not a schema lesson. Mechanics: Structured Outputs.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Deep dive · Live overlaps — pointer only
Logit masks and schema-constrained text: Structured Outputs and Streaming Structured Output. vLLM / PagedAttention is a comparator for paged KV — not a lesson in this cluster. Do not re-teach System One / Jev; that cluster is typed decisions, not generation.
Pitfalls
Given 8×H100, a 3-second TTFT SLO, weekly weight drops, and a JSON-constrained agent loop: which engine, and what symptom would tell you that you picked wrong? Point at the sibling that owns the depth.
Interview Q&A
When would you pick SGLang over TensorRT-LLM?
Answer
SGLang when structured generation, prefix reuse, or rapid model iteration matter. TensorRT-LLM when you need max NVIDIA throughput and can invest in engine builds. Vendor-reported gains must be validated on your workload. Depth: SGLang and TensorRT-LLM.
Explain FSDP vs DDP in one minute.
Answer
DDP replicates the model and splits the batch, all-reducing gradients. FSDP shards parameters, gradients, and optimizer states across ranks, gathering them per layer — larger models, more communication. Depth: Distributed PyTorch.
What symptom suggests missing continuous batching?
Answer
High tail latency and low GPU utilization under concurrent requests, with good single-request performance.
How does KV cache affect serving cost?
Answer
KV grows with context length and concurrency. Poor management causes OOM or recomputation, raising cost per token. Paged KV (vLLM comparator) mitigates fragmentation. Capacity math: playbook.