TensorRT-LLM — Engine Build, In-Flight Batching & Quantization
TensorRT-LLM compiles a model into a serialized TensorRT engine bound to one GPU SKU, one parallelism layout, and one quantization recipe. That buys NVIDIA latency and lock-in. This lesson covers convert → build → run, in-flight batching, FP8/INT4/INT8, multi-GPU topology, and the rebuild tax.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Compiled NVIDIA engine vs a portable runtime
Prefer
TRT-LLM when the fleet is stable NVIDIA
Small, stable model set; known GPU SKU; lowest TTFT or highest tokens/sec per dollar; you can bake quant and TP/PP into CI.
- Fused kernels and IFB plus paged KV are first-class on NVIDIA.
- FP8 on Hopper/Blackwell is the default quality/perf trade (vendor-reported 1.3–2× vs FP16).
- Accept rebuilding on every model, GPU, or driver change.
Alternative
Portable engine when hardware or weights churn
ONNX Runtime, SGLang / vLLM-style Python serving, or llama.cpp when SKUs mix, models change weekly, or the team cannot own a compile pipeline.
- Peak latency will be worse; iteration speed better.
- A daily-changing checkpoint plus TRT-LLM is the wrong-stack smell.
- INT4 on a 1–3B reasoning model is an accuracy cliff, not a model bug.
Convert, build, run
Only the engine runs in production. Shape limits and TP/PP are baked at build time.
- 1
Convert
convert_checkpoint.py reads HF/PyTorch weights and writes a TRT-LLM checkpoint plus config JSON. No GPU code yet. - 2
Build
trtllm-build fuses kernels, assigns TP/PP, bakes max batch/seq/tokens and KV block size, emits .engine files. - 3
Run
C++ runtime or Python LLM / ModelRunner loads engines. Nothing is re-optimized at load. - 4
Rebuild when the tuple changes
New weights, quant, limits, TP/PP, TRT version, driver, or GPU arch invalidates the artifact.
Overview
TensorRT-LLM is NVIDIA’s serving stack for the narrow but expensive case: NVIDIA GPUs, a fixed set of models, and a hard latency or cost-per-token target. Unlike portable runtimes, it compiles a model into a serialized TensorRT engine tuned to one GPU SKU, one parallelism layout, and one quantization recipe.
That buys throughput and time-to-first-token that portable engines rarely match on the same silicon — and it buys lock-in. The artifact you deploy is not a model; it is a compiled plan.
Compile path: PyTorch / HF to TRT engine
Three stages, and only the last one runs in production.
Convert. convert_checkpoint.py (or the model-specific converter) reads a HuggingFace / PyTorch checkpoint and emits a TRT-LLM checkpoint directory: weights in the layout and dtype the builder expects, plus a config JSON describing num_layers, hidden size, TP/PP degrees, and quantization. Still a normal file format; no GPU code is compiled.
Build. trtllm-build consumes the converted checkpoint and emits one or more .engine files. This step picks TensorRT plugins, fuses kernels, assigns layers across GPUs for tensor / pipeline parallelism, and bakes in max batch size, max input length, max sequence length, and paged KV-cache block size. Some quantized paths do calibration or weight-only repacking here. Build is CPU/RAM hungry and can take tens of minutes.
Run. A C++ runtime or the Python LLM / ModelRunner API loads engines and serves generations. Nothing is re-optimized at load; the plan executes.
Flow
- 1
1 HF / PyTorch checkpoint
- next2 convert_checkpoint.py
- 2
2 convert_checkpoint.py
- next3 TRT-LLM ckpt + config
- 3
3 TRT-LLM ckpt + config
- next4 trtllm-build
- 4
4 trtllm-build
- next5 Fuse kernels and plugins
- 5
5 Fuse kernels and plugins
- next6 Quant FP8 / INT4 / INT8
- 6
6 Quant FP8 / INT4 / INT8
- next7 .engine bound to SKU
- 7
7 .engine bound to SKU
- next8 In-flight batching
- 8
8 In-flight batching
- next9 Paged KV blocks
- 9
9 Paged KV blocks
- next10 TensorRT on GPU
- 10
10 TensorRT on GPU
Lesson map
TensorRT-LLM — Engine Build, In-Flight Batching & Quantization
TensorRT-LLM compiles a model into a serialized TensorRT engine bound to one GPU SKU, one parallelism layout, and one quantization recipe. That buys NVIDIA latency and lock-in. This lesson covers convert → build → run, in-flight batching, FP8/INT4/INT8, multi-GPU topology, and the rebuild tax.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 HF / PyTorch checkpoint"] b["2"] c["3 TRT-LLM ckpt + config"] d["4 trtllm-build"] a -->|1 HF / PyTorch checkpoint| b b -->|2 to 3 TRT-LLM ckpt + config| c c -->|3 TRT-LLM ckpt + config| d
In-flight batching
Classic continuous batching queues at the request level: a finished sequence frees its slot and the next request joins. TensorRT-LLM’s in-flight batching (IFB) goes further — generation, context / prefill, and eviction can interleave inside a single forward pass, so a decode step for an old request and the prefill of a new one share the same iteration.
Combined with the paged KV cache (blocks allocated per sequence rather than one contiguous slab), this removes the head-of-line stall where a long prompt blocks everyone’s tokens.
Practically, IFB is configured at build time via max_batch_size, max_num_tokens, max_seq_len, and KV block size. Get those wrong and you either leave GPU idle or crash at runtime with an out-of-blocks error. There is no dynamic growth: an engine built for 8192 context will refuse a 12288-token prompt.
FP8 / INT4 / INT8
- FP8 (E4M3) — default on Hopper / Blackwell-class parts. Weights and activations in 8-bit float; needs scaling factors. Near-BF16 quality on most LLMs; large memory and bandwidth win. Vendor-reported speedups are commonly 1.3–2× over FP16 on supported hardware.
- INT8 with SmoothQuant — 8-bit weights and activations, with an activation-smoothing transform that moves quantization difficulty into weights. Works on older GPUs that lack FP8.
- INT4 weight-only (AWQ / GPTQ-style) — 4-bit weights, 16-bit compute. Best memory savings; quality risk rises, especially on small models and reasoning-heavy tasks.
All vendor latency and throughput figures here are vendor-reported and hardware-specific. Benchmark your own SKU, sequence mix, and concurrency before trusting them.
Multi-GPU
Parallelism is chosen at build time, not runtime. Tensor parallelism (TP) splits each layer’s matmuls across GPUs with all-reduce between them. Pipeline parallelism (PP) splits layers into stages. Expert parallelism applies to MoE models.
Because the engine’s topology is compiled in, a 4-GPU TP engine is useless on 2 GPUs. TP wants fast intra-node NVLink; PP tolerates slower inter-node links but adds bubble latency. Multi-GPU builds are also where quantization interacts badly: per-channel scales must be consistent across ranks or numerics drift.
Training-style DDP/FSDP is the Distributed PyTorch lesson — serving TP/PP here is a compiled layout, not a gradient sync.
When NVIDIA-optimized latency wins
Choose TensorRT-LLM when: you serve a small, stable model fleet on known NVIDIA GPUs; you need the lowest TTFT or highest tokens/sec per dollar; you can bake quantization and parallelism into CI; and you accept rebuilding on every model, GPU, or driver change.
Choose a portable engine (ONNX Runtime, vLLM / SGLang-style Python serving, llama.cpp) when: hardware is heterogeneous or may change; models churn weekly; you need CPU or non-NVIDIA accelerators; or the team cannot own a compile pipeline. Peak latency will be worse, iteration speed better. Portable path depth: ONNX Runtime, SGLang.
Engine rebuild cost
An engine is invalidated by: a new model version or weight edit, a different quant recipe, a change in max batch / seq / token limits, a different TP/PP degree, a new TensorRT or TRT-LLM version, and often a driver or GPU-arch change.
Rebuild is wall-clock minutes to tens of minutes plus validation, and it is per (model, GPU, layout) tuple. If your fleet mixes SKUs, multiply. There is no single TRT-LLM engine.
Comparative pros / cons + wrong-stack
Pros: best-in-class NVIDIA latency and memory efficiency; mature IFB and paged KV; first-class FP8 / INT4; strong multi-GPU support; production-proven at NVIDIA scale.
Cons: compile step in the critical path; hard binding to GPU SKU and driver; large container images; Python-first APIs with a C++ runtime that is harder to debug; quantization quality cliffs on small models.
Wrong-stack smells: building TRT-LLM engines for a model that changes daily; running it on AMD / CPU / Apple silicon; picking INT4 for a 1–3B reasoning model and blaming the model; using a 70B-class TP-8 engine on a single-GPU dev box; treating a nightly rebuild as “just CI”.
Flow
- 1
1 Export HF / PyTorch
- next2 TRT-LLM build
- 2
2 TRT-LLM build
- next3 Engine FP8 / INT
- 3
3 Engine FP8 / INT
- next4 In-flight batching
- 4
4 In-flight batching
- next5 Tokens to client
- 5
5 Tokens to client
Sandbox: mock engine build (Python)
Deterministic stand-in so build logic can be unit-tested without CUDA.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Short TypeScript client
API shape only. No GPU credentials.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Deep dive · Portable stacks — pointer only
ONNX export and execution providers: ONNX & ONNX Runtime. Prefix-heavy Python serving: SGLang. vLLM is named as a portable alternative, not taught. Metrics and rebuild tax in the playbook.
Pitfalls
List five changes that force a rebuild. Which of them can you pin in CI, and which come from a GPU SKU mix you should have refused?
Interview Q&A
Why does TensorRT-LLM need an engine build step that SGLang / vLLM-class servers often skip?
Answer
TRT-LLM compiles the model into a TensorRT engine with fused kernels, static/dynamic shape profiles, and quantization baked in. That front-loads optimization for NVIDIA latency; the cost is rebuilds when weights, precision, max seq/batch, or CUDA/TRT versions change.
What is in-flight batching, and how does it differ from a naive static batch?
Answer
In-flight (continuous-style) batching admits new requests as slots free during decode instead of waiting for an entire fixed batch to finish. TRT-LLM can also interleave prefill and decode inside one GPU iteration. Improves GPU utilization and p99 under mixed prompt lengths.
When would you pick TRT-LLM over a portable ONNX Runtime or SGLang stack?
Answer
When you are NVIDIA-locked, need vendor-optimized lowest latency after measure-twice benchmarks on your model/HW, and can afford engine CI. Prefer portable stacks when multi-vendor CPUs/GPUs, fast model iteration, or ops cannot own rebuild pipelines.
Name two failure modes unique to compiled engines.
Answer
(1) Stale engines after CUDA/driver/TRT upgrades → crashes or silent perf cliffs. (2) Shape-profile mismatch (longer context / larger batch than built) → runtime errors or fallback paths. Always version engines beside model + CUDA pins.
Go Deeper
Official NVIDIA TensorRT-LLM documentation and repository only: