LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offs
RTN and per-tensor/channel/group scales, outliers, WxAy notation; GPTQ vs AWQ (W4A16), SmoothQuant (W8A8), FP8 E4M3; calibration, accuracy checks, when it helps TTFT vs only TPOT; runnable quant basics, GPTQ/AWQ toy, format math.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does W4A16 mean?
Answer
4-bit weights dequantised to 16-bit inside the kernel; the matmul runs in 16-bit.
L2
Why do per-tensor scales fail?
Answer
A few outlier channels stretch the scale so most values collapse into a few bins.
L3
What does GPTQ optimise?
Answer
Layer output error, compensating each column's error in the remaining weights with inverse-Hessian information.
L4
What does AWQ do?
Answer
Scales up the weights of the few salient input channels before quantization.
L5
What problem does SmoothQuant solve?
Answer
Activation outliers that make INT8 activations lossy.
L6
Why is FP8 easier than INT8?
Answer
Its exponent bits give a non-uniform grid with wide dynamic range.
L7
How do you decide a quantized model is safe?
Answer
A quality budget on task evals, a benchmark at production batch sizes, then a canary.
Failure modes
INT4 for a TTFT problem
Prefill still runs 16-bit math, so TTFT barely moves.
Wrong calibration data
Salience and Hessian estimates miss real traffic.
INT4 model with no KV room
It fits on one GPU but serves batch 1.
Misconceptions
Smaller weight error means better quantization.
GPTQ's weights were further from the originals yet its outputs were closer.
Perplexity is a sufficient check.
Perplexity can move 1 percent while a reasoning benchmark moves 5 points.
Every engine has fast kernels for every format.
Check the support matrix per GPU.
Interviewer traps
Comparing formats at different batch sizes or context lengths.
Benchmark at production batch sizes and context lengths.
Quantizing a fine-tuned model with pre-fine-tune calibration data.
Calibrate on data matching the fine-tuned model's traffic.
Design scenario
Same prompt for every reader.
Requirements
Lower cost per token, keep task accuracy within budget, and know whether TTFT or only TPOT improves.
Failure assumptions
- Calibration data is generic web text.
- The code model is sensitive to INT4.
- LoRA adapters were trained on the BF16 base.
Constraints
- Hopper for one fleet, Ampere for the other.
- A fixed accuracy budget per release.
Prompt
Choose a quantization format for a 70B chat model on H100s and an 8B code model on A100s.
API
Which formats does the engine support per GPU, and how are quantized checkpoints versioned?
Data
Which calibration sets and eval slices (long-context, reasoning, safety) are used?
Architecture
Which format per fleet, and how is the canary against BF16 run?
Overview
Decode is memory-bandwidth-bound: every output token reads every weight (and the KV cache) from HBM. Storing weights in fewer bits therefore buys speed and memory at once. That is why weight quantization is the first cost lever most teams pull. The engineering question is not "quantize or not" but which format, at which granularity, calibrated how, verified on what. This page compares the main families: round-to-nearest (RTN) with per-tensor, per-channel and per-group scales; weight-only INT4 (W4A16) via GPTQ (second-order error compensation) and AWQ (activation-aware scaling of salient channels); W8A8 INT8 made possible by SmoothQuant; and FP8 (E4M3) on Hopper-class and newer GPUs. It covers how each works, what it costs in accuracy, and when it helps TTFT vs only TPOT. The TensorRT-LLM page already shows how to build quantized engines; this page is the mechanics and the trade-offs underneath.
The vocabulary
- WxAy: weights in x bits, activations (and so the matmul) in y bits. W4A16 = 4-bit weights dequantised to 16-bit inside the kernel; W8A8 = both 8-bit, so the matmul itself runs on 8-bit tensor cores.
- Scale and zero-point:
q = round(w / s) + z;w ~ s * (q - z). Symmetric quantization drops z. - Granularity: one scale per tensor, per output channel (row), or per group of g consecutive weights (g = 128 is the common default for INT4).
- Calibration: running a few hundred sample prompts to collect activation statistics. GPTQ, AWQ and SmoothQuant need it; RTN does not.
- KV-cache quantization: a separate choice (FP8 KV is the common one); it attacks the other big source of decode bytes.
The formats, side by side
| Format | Bytes vs BF16 | Speeds prefill? | Speeds decode? | Typical accuracy impact | Hardware | Good for |
|---|---|---|---|---|---|---|
| BF16 | 1x | - | - | Reference | Any modern GPU | Quality baseline |
| FP8 W8A8 (+ FP8 KV) | 0.5x | Yes | Yes | Near-lossless on most models | Hopper+ | Default production choice on H100/H200/B200 |
| INT8 W8A8 (SmoothQuant) | 0.5x | Yes | Yes | Small, needs smoothing | Ampere+ | A100-class fleets without FP8 |
| INT8 W8A16 | 0.5x | No | Yes | Very small | Any | Simple memory saving |
| INT4 W4A16 g128 (GPTQ/AWQ) | about 0.27x | No (16-bit math) | Yes, most | Noticeable on small models, long context, math and code | Any with kernels (Marlin, etc.) | Fit big models on fewer GPUs, low-batch latency, edge |
| 4-bit float (NVFP4/MXFP4) W4A4 | about 0.28x | Yes | Yes | Depends on recipe; verify per model | Blackwell+ | Throughput on newest hardware |
FP8 W8A8 or INT4 W4A16?
Prefer
FP8 W8A8 where hardware supports it
Half the bytes, 8-bit tensor-core math.
- Speeds prefill and decode.
- 70B fits in 70 GB on one GPU in the format math.
- Near-lossless on most models with simple calibration.
Alternative
INT4 W4A16 g128 (GPTQ/AWQ)
About 0.27x the bytes, 16-bit math.
- 8B batch-1 ceiling of 812 tok/s vs 419 for FP8.
- No prefill speed-up.
- Noticeable accuracy loss on small models, long context, math and code.
Quantize and ship
Diagram 1 condensed: BF16 checkpoint to quantized deployment, with the failure path.
- 1
Pick a format
By hardware and whether TTFT or TPOT is the problem. - 2
Calibrate
A few hundred prompts that match real traffic. - 3
Quantize
RTN, GPTQ, AWQ, SmoothQuant or FP8. - 4
Evaluate
Task, long-context, reasoning and safety slices, then a latency benchmark. - 5
Quality budget exceeded
Change format or granularity; do not ship on perplexity alone.
// Weight format -> memory, GPU count, and the batch-1 decode speed ceiling.
// Decode at small batch is memory-bandwidth bound: each token must stream every weight once,
// so tokens/s per sequence <= bandwidth / weight bytes. Example GPU: 80 GB, 3.35 TB/s (H100 SXM-class).
// Simplified: ignores KV reads, activations, kernel efficiency and multi-GPU communication.
type Fmt = { name: string; bitsPerWeight: number; compute: string; note: string };
const formats: Fmt[] = [
{ name: "BF16", bitsPerWeight: 16, compute: "BF16 tensor cores", note: "reference quality" },
{ name: "FP8 W8A8", bitsPerWeight: 8, compute: "FP8 tensor cores (Hopper+)", note: "near-lossless on most LLMs" },
{ name: "INT8 W8A8 (SmoothQuant)", bitsPerWeight: 8, compute: "INT8 tensor cores (Ampere+)", note: "needs outlier smoothing" },
{ name: "INT4 W4A16 g128 (AWQ/GPTQ)", bitsPerWeight: 4 + 16 / 128, compute: "dequant to FP16 in kernel", note: "small models & reasoning lose most" },
];
const models: Array<[string, number]> = [["8B", 8e9], ["70B", 70e9], ["405B", 405e9]];
const HBM = 80e9, BW = 3.35e12, USABLE = 0.9;
console.log("format model weights minGPUs ceiling@minGPUs ceiling@8GPUs (tok/s, batch 1)");
for (const f of formats) {
for (const [m, p] of models) {
const bytes = (p * f.bitsPerWeight) / 8;
const gpus = Math.ceil(bytes / (HBM * USABLE)); // weights only; KV needs more
const ceiling = (BW * gpus) / bytes; // tensor-parallel shards stream in parallel
const at8 = gpus <= 8 ? ((BW * 8) / bytes).toFixed(0) : "n/a";
console.log(`${f.name.padEnd(30)}${m.padEnd(6)}${(bytes / 1e9).toFixed(0).padStart(6)} GB${String(gpus).padStart(7)}${ceiling.toFixed(0).padStart(15)}${at8.padStart(15)}`);
}
}
console.log("\nformat notes:");
formats.forEach((f) => console.log(` ${f.name.padEnd(28)} compute: ${f.compute.padEnd(28)} ${f.note}`));
console.log("\nTakeaways: on a fixed node, halving bytes doubles the batch-1 ceiling; or keep speed and halve GPU count.");
console.log("W4A16 helps decode (memory-bound) but not prefill (compute-bound, still 16-bit math).");
console.log("Failure path: an INT4 model that 'fits on 1 GPU' with no room left for KV cache serves batch 1 only.");Output:
format model weights minGPUs ceiling@minGPUs ceiling@8GPUs (tok/s, batch 1)
BF16 8B 16 GB 1 209 1675
BF16 70B 140 GB 2 48 191
BF16 405B 810 GB 12 50 n/a
FP8 W8A8 8B 8 GB 1 419 3350
FP8 W8A8 70B 70 GB 1 48 383
FP8 W8A8 405B 405 GB 6 50 66
INT8 W8A8 (SmoothQuant) 8B 8 GB 1 419 3350
INT8 W8A8 (SmoothQuant) 70B 70 GB 1 48 383
INT8 W8A8 (SmoothQuant) 405B 405 GB 6 50 66
INT4 W4A16 g128 (AWQ/GPTQ) 8B 4 GB 1 812 6497
INT4 W4A16 g128 (AWQ/GPTQ) 70B 36 GB 1 93 743
INT4 W4A16 g128 (AWQ/GPTQ) 405B 209 GB 3 48 128
format notes:
BF16 compute: BF16 tensor cores reference quality
FP8 W8A8 compute: FP8 tensor cores (Hopper+) near-lossless on most LLMs
INT8 W8A8 (SmoothQuant) compute: INT8 tensor cores (Ampere+) needs outlier smoothing
INT4 W4A16 g128 (AWQ/GPTQ) compute: dequant to FP16 in kernel small models & reasoning lose most
Takeaways: on a fixed node, halving bytes doubles the batch-1 ceiling; or keep speed and halve GPU count.
W4A16 helps decode (memory-bound) but not prefill (compute-bound, still 16-bit math).
Failure path: an INT4 model that 'fits on 1 GPU' with no room left for KV cache serves batch 1 only.Expectedformat model weights minGPUs ceiling@minGPUs ceiling@8GPUs (tok/s, batch 1) BF16 8B 16 GB 1 209 1675 BF16 70B 140 GB 2 48 191 BF16 405B 810 GB 12 50 n/a FP8 W8A8 8B 8 GB 1 419 3350 FP8 W8A8 70B 70 GB 1 48 383 FP8 W8A8 405B 405 GB 6 50 66 INT8 W8A8 (SmoothQuant) 8B 8 GB 1 419 3350 INT8 W8A8 (SmoothQuant) 70B 70 GB 1 48 383 INT8 W8A8 (SmoothQuant) 405B 405 GB 6 50 66 INT4 W4A16 g128 (AWQ/GPTQ) 8B 4 GB 1 812 6497 INT4 W4A16 g128 (AWQ/GPTQ) 70B 36 GB 1 93 743 INT4 W4A16 g128 (AWQ/GPTQ) 405B 209 GB 3 48 128 format notes: BF16 compute: BF16 tensor cores reference quality FP8 W8A8 compute: FP8 tensor cores (Hopper+) near-lossless on most LLMs INT8 W8A8 (SmoothQuant) compute: INT8 tensor cores (Ampere+) needs outlier smoothing INT4 W4A16 g128 (AWQ/GPTQ) compute: dequant to FP16 in kernel small models & reasoning lose most Takeaways: on a fixed node, halving bytes doubles the batch-1 ceiling; or keep speed and halve GPU count. W4A16 helps decode (memory-bound) but not prefill (compute-bound, still 16-bit math). Failure path: an INT4 model that 'fits on 1 GPU' with no room left for KV cache serves batch 1 only.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Granularity and outliers: why naive quantization fails
LLM activations have a few outlier channels with magnitudes 10-100x the rest (documented in LLM.int8 and SmoothQuant). Weights have milder outliers. A single per-tensor scale must stretch to cover the largest value, so most values collapse into a handful of bins.
"""Quantization mechanics on synthetic LLM-like tensors (numpy only).
1) Granularity: per-tensor vs per-channel INT8 vs group-wise INT4 (g=128) for WEIGHTS.
2) Activation outliers: a few hidden channels are ~50x larger (seen in real LLMs), which
wrecks per-tensor INT8 activation quantization.
3) SmoothQuant: divide activations by s_j and multiply weights by s_j (mathematically identical
product) so the outlier difficulty migrates from activations into weights.
"""
import numpy as np
rng = np.random.default_rng(0)
def qdq(x, bits, axis=None, group=None, sym=True):
"""Quantize then dequantize. axis: reduce over this axis for scales; group: group size on last axis."""
qmax = 2 ** (bits - 1) - 1
if group:
shp = x.shape
x = x.reshape(shp[0], -1, group)
s = np.abs(x).max(axis=-1, keepdims=True) / qmax
return (np.clip(np.round(x / s), -qmax - 1, qmax) * s).reshape(shp)
s = np.abs(x).max(axis=axis, keepdims=axis is not None) / qmax
return np.clip(np.round(x / s), -qmax - 1, qmax) * s
def rel(a, b): return np.linalg.norm(a - b) / np.linalg.norm(a)
d_in, d_out, tokens = 1024, 1024, 256
W = rng.normal(0, 0.02, (d_out, d_in))
W[:8] *= 6 # a few output rows with larger weights (uneven ranges)
print("WEIGHT error (relative Frobenius):")
print(f" INT8 per-tensor {rel(W, qdq(W, 8)):.4f}")
print(f" INT8 per-channel {rel(W, qdq(W, 8, axis=1)):.4f} (one scale per output row)")
print(f" INT4 per-channel {rel(W, qdq(W, 4, axis=1)):.4f}")
print(f" INT4 group g=128 {rel(W, qdq(W, 4, group=128)):.4f} (one scale per 128 weights, AWQ/GPTQ default)")
X = rng.normal(0, 1, (tokens, d_in))
outlier_ch = rng.choice(d_in, 6, replace=False)
X[:, outlier_ch] *= 50 # systematic activation outlier channels
Y = X @ W.T
print("\nOUTPUT error of X @ W^T with W8A8 (both INT8):")
Yq = qdq(X, 8, axis=1) @ qdq(W, 8, axis=1).T # per-token activation scales, per-channel weight scales
print(f" naive W8A8 {rel(Y, Yq):.4f}")
alpha = 0.5 # migration strength from the SmoothQuant paper
s = np.abs(X).max(axis=0) ** alpha / np.abs(W).max(axis=0) ** (1 - alpha)
Xs, Ws = X / s, W * s # (X/s) @ (W*s)^T == X @ W^T exactly
assert np.allclose(Xs @ Ws.T, Y)
Yq_s = qdq(Xs, 8, axis=1) @ qdq(Ws, 8, axis=1).T
print(f" SmoothQuant W8A8 a=0.5 {rel(Y, Yq_s):.4f}")
print(f" weight-only W8A16 {rel(Y, X @ qdq(W, 8, axis=1).T):.4f} (activations stay 16-bit)")
print("\nFailure path: per-tensor scales on outlier-heavy tensors push most values into a few")
print("quantization bins. Fixes: finer scales (per-channel / per-group), smoothing (SmoothQuant),")
print("weight-only quantization, or FP8 whose exponent bits tolerate wide ranges better than INT8.")Output:
WEIGHT error (relative Frobenius):
INT8 per-tensor 0.0471
INT8 per-channel 0.0078 (one scale per output row)
INT4 per-channel 0.1414
INT4 group g=128 0.1171 (one scale per 128 weights, AWQ/GPTQ default)
OUTPUT error of X @ W^T with W8A8 (both INT8):
naive W8A8 0.0502
SmoothQuant W8A8 a=0.5 0.0114
weight-only W8A16 0.0079 (activations stay 16-bit)
Failure path: per-tensor scales on outlier-heavy tensors push most values into a few
quantization bins. Fixes: finer scales (per-channel / per-group), smoothing (SmoothQuant),
weight-only quantization, or FP8 whose exponent bits tolerate wide ranges better than INT8.Reading the output: per-channel scales cut INT8 weight error by about 6x; INT4 needs groups; and SmoothQuant makes W8A8 viable by migrating activation outliers into the weights with a per-channel factor s_j = max|X_j|^alpha / max|W_j|^(1-alpha) (alpha around 0.5). The math is unchanged because X W^T = (X / s)(W * s)^T, but both factors become easier to quantize.
GPTQ vs AWQ: two ways to choose what error to keep
Both target W4A16 (or W3) with groups, both use calibration data, and both beat RTN. They differ in what they optimise.
| RTN | GPTQ | AWQ | |
|---|---|---|---|
| Idea | Round each weight to the nearest level | Quantize columns one at a time; update the remaining columns to cancel the error, using the inverse Hessian of layer inputs (Optimal Brain Surgeon lineage) | Find the ~1 percent of input channels with large activations; scale those weights up before quantization (and inputs down) so they lose less precision |
| Optimises | Weight error | Layer output error on calibration data | Output error via per-channel scaling search |
| Calibration | None | Yes (Hessian), sensitive to data | Yes (activation magnitudes), less prone to overfitting calibration set per the AWQ paper |
| Quantization time | Seconds | Minutes to hours for large models | Minutes |
| Kernel at serve time | Same W4A16 kernels | Same | Same |
"""Why calibration-based 4-bit methods beat round-to-nearest (RTN): toy GPTQ and AWQ (numpy).
Both use a small CALIBRATION set of real activations X. Neither retrains the model.
GPTQ: quantize weights one input column at a time; push each rounding error onto the
not-yet-quantized columns using the inverse Hessian H^-1, H = X^T X (OBS-style update),
so the layer OUTPUT X @ w is preserved, not the weights themselves.
AWQ : find the few input channels with large activations ("salient"), scale those weight
columns UP by s (and the activations down by s) before quantizing, which shrinks
their relative rounding error. Search the scaling exponent on calibration data.
"""
import numpy as np
rng = np.random.default_rng(1)
d, rows, n = 128, 64, 1024
BITS = 3 # 3-bit makes the differences easy to see
QMAX = 2 ** (BITS - 1) - 1
mix = rng.normal(0, 1, (d, d)) / np.sqrt(d)
chan_scale = np.ones(d); chan_scale[rng.choice(d, 4, replace=False)] = 8.0 # salient channels
def acts(m): # correlated activations with a few large channels
return (rng.normal(0, 1, (m, d)) @ mix + 0.3 * rng.normal(0, 1, (m, d))) * chan_scale
Xcal, Xtest = acts(n), acts(n)
W = rng.normal(0, 0.05, (rows, d))
def rtn(w, scale):
return np.clip(np.round(w / scale), -QMAX - 1, QMAX) * scale
def gptq(W, X, damp=0.01):
H = X.T @ X
H += damp * np.mean(np.diag(H)) * np.eye(d) # dampening for numerical stability
Q = np.zeros_like(W)
for r in range(W.shape[0]):
w = W[r].copy(); Hinv = np.linalg.inv(H)
scale = np.abs(W[r]).max() / QMAX # per-row scale fixed up front
for i in range(d):
Q[r, i] = rtn(w[i], scale)
err = (w[i] - Q[r, i]) / Hinv[i, i]
w[i + 1:] -= err * Hinv[i, i + 1:] # compensate on remaining columns
Hinv -= np.outer(Hinv[:, i], Hinv[i, :]) / Hinv[i, i] # remove column i
return Q
def awq(W, X, alphas=np.linspace(0, 1, 11)):
a_mag = np.abs(X).mean(axis=0)
best = None
for a in alphas:
s = a_mag ** a; s /= np.sqrt(s.max() * s.min()) # normalise around 1
Ws = W * s
Q = np.stack([rtn(row, np.abs(row).max() / QMAX) for row in Ws]) / s # fold 1/s back
err = np.linalg.norm(X @ W.T - X @ Q.T)
if best is None or err < best[0]:
best = (err, a, Q)
return best[2], best[1]
def out_err(Q, X): return np.linalg.norm(X @ W.T - X @ Q.T) / np.linalg.norm(X @ W.T)
Q_rtn = np.stack([rtn(row, np.abs(row).max() / QMAX) for row in W])
Q_gptq = gptq(W, Xcal)
Q_awq, a = awq(W, Xcal)
print(f"{BITS}-bit per-row weights, relative OUTPUT error on held-out activations:")
print(f" RTN {out_err(Q_rtn, Xtest):.4f}")
print(f" AWQ (alpha={a:.1f}) {out_err(Q_awq, Xtest):.4f}")
print(f" GPTQ {out_err(Q_gptq, Xtest):.4f}")
print(f"weight-space error: RTN {np.linalg.norm(W-Q_rtn)/np.linalg.norm(W):.3f} GPTQ {np.linalg.norm(W-Q_gptq)/np.linalg.norm(W):.3f}")
print("GPTQ's weights are FURTHER from the originals yet its outputs are closer: it optimises what matters.")
print("Failure path: calibrate on the wrong distribution (e.g. English-only for a code model) and the")
print("Hessian/salience estimates are wrong for real traffic; re-run evals on your own domain.")Output:
3-bit per-row weights, relative OUTPUT error on held-out activations:
RTN 0.2684
AWQ (alpha=0.3) 0.2128
GPTQ 0.2272
weight-space error: RTN 0.268 GPTQ 0.409
GPTQ's weights are FURTHER from the originals yet its outputs are closer: it optimises what matters.
Failure path: calibrate on the wrong distribution (e.g. English-only for a code model) and the
Hessian/salience estimates are wrong for real traffic; re-run evals on your own domain.Note the counter-intuitive line: GPTQ's weights end up further from the originals than RTN's, yet its outputs are closer. Quantization quality is about the function, not the weights.
FP8: the 2026 default where hardware supports it
FP8 E4M3 (4 exponent bits, 3 mantissa bits) has a floating-point grid: dense near zero, sparse at the extremes. That tolerates outliers far better than INT8's uniform grid, so per-tensor or per-channel FP8 W8A8 is near-lossless on most LLMs with simple calibration, and it runs on FP8 tensor cores (Hopper, Blackwell and newer), so it speeds up prefill as well as decode. Blackwell-class GPUs also add 4-bit floating formats (MXFP4/NVFP4) with fine-grained block scales, which extend the same idea to 4 bits.
The quantize-and-ship workflow
Diagram 1: from BF16 checkpoint to a quantized deployment, with the failure path.
Decisions
- 1
Step 1: BF16 checkpoint plus baseline evals (task set, long context, reasoning, safety)
- nextStep 2: pick format from hardware and bottleneck (FP8 on Hopper+, INT4 to fit, INT8 SmoothQuant on Ampere)
- 2
Step 2: pick format from hardware and bottleneck (FP8 on Hopper+, INT4 to fit, INT8 SmoothQuant on Ampere)
- nextStep 3: collect calibration prompts drawn from real traffic and domain
- 3
Step 3: collect calibration prompts drawn from real traffic and domain
- nextStep 4: quantize (RTN or FP8 scales, AWQ search, or GPTQ Hessian updates) with groups if 4-bit
- 4
Step 4: quantize (RTN or FP8 scales, AWQ search, or GPTQ Hessian updates) with groups if 4-bit
- nextStep 5: re-run the same evals plus a perplexity sanity check
- 5
Step 5: re-run the same evals plus a perplexity sanity check
- nextStep 6: within the agreed quality budget?
- ?
Step 6: within the agreed quality budget?
- nextStep 7: benchmark TTFT, TPOT, throughput on the serving engine at target batch
- nextFailure path: math or long-context scores drop, outputs repeat or degrade
- 7
Step 7: benchmark TTFT, TPOT, throughput on the serving engine at target batch
- nextStep 8: canary, compare online metrics with the BF16 deployment
- 8
Step 8: canary, compare online metrics with the BF16 deployment
- 9
Failure path: math or long-context scores drop, outputs repeat or degrade
- nextUse finer groups, better calibration data, keep sensitive layers (embeddings, lm_head) in 16-bit, or step up to 8-bit
- 10
Use finer groups, better calibration data, keep sensitive layers (embeddings, lm_head) in 16-bit, or step up to 8-bit
- nextStep 4: quantize (RTN or FP8 scales, AWQ search, or GPTQ Hessian updates) with groups if 4-bit
Lesson map
LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offs
RTN and per-tensor/channel/group scales, outliers, WxAy notation; GPTQ vs AWQ (W4A16), SmoothQuant (W8A8), FP8 E4M3; calibration, accuracy checks, when it helps TTFT vs only TPOT; runnable quant basics, GPTQ/AWQ toy, format math.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: BF16 checkpoint plus baseline evals (task set, long context, reasoning, safety)"] s2["Step 2: pick format from hardware and bottleneck (FP8 on Hopper+, INT4 to fit, INT8 SmoothQuant on Ampere)"] s3["Step 3: collect calibration prompts drawn from real traffic and domain"] s4["Step 4: quantize (RTN or FP8 scales, AWQ search, or GPTQ Hessian updates) with groups if 4-bit"] s5["Step 5: re-run the same evals plus a perplexity sanity check"] s6["Step 6: within the agreed quality budget?"] s7["Step 7: benchmark TTFT, TPOT, throughput on the serving engine at target batch"] s8["Step 8: canary, compare online metrics with the BF16 deployment"] f1["Failure path: math or long-context scores drop, outputs repeat or degrade"] f2["Use finer groups, better calibration data, keep sensitive layers (embeddings, lm_head) in 16-bit, or step up to 8-bit"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s4 -->|continues| s5 s5 -->|continues| s6 s6 -->|continues| s7 s7 -->|continues| s8 s6 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s4
Decision chart: which format?
Decisions
- 1
D1
- nextStep 2: does an 8-bit version fit with KV room?
- nextStep 3: GPU has FP8 tensor cores (Hopper or newer)?
- ?
Step 2: does an 8-bit version fit with KV room?
- nextStep 3: GPU has FP8 tensor cores (Hopper or newer)?
- nextINT4 W4A16 (AWQ or GPTQ, g128) or 4-bit float on Blackwell; evaluate carefully
- ?
Step 3: GPU has FP8 tensor cores (Hopper or newer)?
- nextFP8 W8A8 plus FP8 KV cache - default
- nextStep 4: is prefill compute a bottleneck?
- 4
INT4 W4A16 (AWQ or GPTQ, g128) or 4-bit float on Blackwell; evaluate carefully
- nextAlways: task evals, long context and reasoning slices before shipping
- 5
FP8 W8A8 plus FP8 KV cache - default
- nextAlways: task evals, long context and reasoning slices before shipping
- ?
Step 4: is prefill compute a bottleneck?
- nextINT8 W8A8 with SmoothQuant
- nextWeight-only INT8 or INT4, plus KV quantization if memory-bound
- 7
INT8 W8A8 with SmoothQuant
- nextAlways: task evals, long context and reasoning slices before shipping
- 8
Weight-only INT8 or INT4, plus KV quantization if memory-bound
- nextAlways: task evals, long context and reasoning slices before shipping
- 9
Always: task evals, long context and reasoning slices before shipping
What happens if you choose otherwise
- INT4 for a TTFT problem: prefill still runs 16-bit math; TTFT barely moves while accuracy risk rises.
- Per-tensor INT8 activations without smoothing: outlier channels wreck accuracy; it looks like a broken model.
- Calibrate on generic web text for a code or multilingual model: the salience and Hessian estimates miss your traffic and quality drops where it matters.
- Squeeze a 70B model onto one GPU at INT4 with no KV room: it "fits" but serves batch 1; two GPUs at FP8 may give more throughput per dollar.
- Ship on perplexity alone: perplexity can move by 1 percent while a reasoning benchmark moves by 5 points.
Pitfalls
- Quantizing a fine-tuned model with calibration data from before the fine-tune.
- Forgetting that LoRA adapters trained on a BF16 base will see a slightly different quantized base at serve time; evaluate the combination.
- Comparing throughput across formats at different batch sizes or context lengths.
- Assuming every engine has fast kernels for every format on every GPU; check the support matrix.
Interview Q&A
Why does weight quantization speed up decode but W4A16 not prefill?
Answer
Decode is memory-bound, so fewer bytes per weight means faster steps. Prefill is compute-bound, and W4A16 still does 16-bit math after dequantising, so it gains little. W8A8 formats (FP8, INT8) speed up both because the matmul runs on 8-bit tensor cores.
GPTQ vs AWQ?
Answer
Both are calibration-based, group-wise INT4 methods. GPTQ quantizes columns sequentially and compensates the error in the remaining weights using second-order (Hessian) information; AWQ protects the small set of salient input channels by scaling them before quantization. Both minimise output error rather than weight error.
What problem does SmoothQuant solve?
Answer
Activation outlier channels make INT8 activation quantization lossy. SmoothQuant divides activations by a per-channel factor and multiplies the weights by it, an exact transformation that moves difficulty from activations to weights so both quantize well.
Why is FP8 easier than INT8?
Answer
Its exponent bits give a non-uniform grid with wide dynamic range, so outliers do not force a huge scale that starves the small values.
How do you decide a quantized model is safe to ship?
Answer
A pre-agreed quality budget on task evals including long-context, reasoning and safety slices, then a latency and throughput benchmark at production batch sizes, then a canary against the BF16 deployment.
Why does a per-channel scale help so much for INT8 weights?
Answer
One scale per output row stops a single outlier from stretching the whole tensor; in the demo INT8 weight error fell from 0.0471 to 0.0078.
What is calibration and which methods need it?
Answer
Running a few hundred sample prompts to collect activation statistics; GPTQ, AWQ and SmoothQuant need it, RTN does not.
What do Blackwell's 4-bit float formats add?
Answer
MXFP4 and NVFP4 with fine-grained block scales extend the FP8 idea to 4 bits, speeding both prefill and decode on Blackwell and newer.
Check yourself
For one model you serve, write its BF16 weight bytes, pick two candidate formats from the table, and list the eval slices and the latency benchmark you would run before a canary.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: TensorRT-LLM — Engine Build, In-Flight Batching & Quantization, Prefill vs Decode & KV Cache Mechanics, Vector Indexes — HNSW, IVF & Product Quantization Tradeoffs, LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter Merging.
Go Deeper
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- SmoothQuant: Accurate and Efficient Post-Training Quantization for LLMs
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- FP8 Formats for Deep Learning
- vLLM quantization docs (GitHub)
- NVIDIA TensorRT Model Optimizer
- Hugging Face Transformers quantization overview