LLM Inference in Production - Quantization, Prompt Caching, Multi-LoRA Serving & Capacity Planning
Hub: fleet-level serving levers beyond the existing runtime pages; which lever (quantization, prefix caching, multi-LoRA, batching/capacity) moves TTFT, TPOT, memory or $/1M; runnable lever simulation + symptom-to-lever picker; triage chart.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What do the runtime pages cover vs this cluster?
Answer
How one request runs vs running a fleet for money and SLOs.
L2
Which lever helps TTFT?
Answer
Prefix caching, queueing control, and chunked prefill or P/D disaggregation.
L3
Which lever helps TPOT?
Answer
Weight and KV quantization, batch size limits and speculative decoding.
L4
Why can FP8 KV matter as much as FP8 weights?
Answer
With long contexts and large batches, KV bytes read per decode step exceed weight bytes.
L5
Does W4A16 help TTFT?
Answer
Not much: prefill is compute-bound and still runs 16-bit math.
L6
When is multi-LoRA worth it?
Answer
Several fine-tunes of the same base with uneven traffic.
L7
Why do levers interact?
Answer
Quantization frees memory that caching and batching then consume.
Failure modes
Quantizing first for a TTFT problem
Weight quantization barely changes queueing or shared-prefix prefill.
Timestamp at the top of the prompt
Prefix cache hit rate stays near zero.
Capacity planned from averages
The fleet looks 60 percent utilised while p95 TTFT breaks at peak.
Misconceptions
Weights are always the dominant bytes per decode step.
With long prompts the KV cache dominates.
Per-request TPOT and fleet tokens per second are the same thing.
One is a request's decode speed, the other is throughput.
Perplexity is enough to approve a quantized model.
Run task evals, long-context and reasoning sets.
Interviewer traps
Serving 40 customer fine-tunes as 40 deployments.
Serve LoRA adapters on one base.
Benchmarking caching with synthetic prompts that share nothing or everything.
Use production-shaped prompts.
Design scenario
Same prompt for every reader.
Requirements
Cut cost per million tokens, keep p95 TTFT within SLO, and consolidate fine-tunes.
Failure assumptions
- A timestamp sits at the top of the system prompt.
- Some replicas run out of KV memory at peak.
- Autoscaling uses average GPU utilisation.
Constraints
- H100-class GPUs.
- No change to the runtime engine this quarter.
Prompt
You inherit an over-budget LLM fleet serving a RAG chatbot with a long system prompt and 40 customer fine-tunes. Prioritise the levers.
API
Which metrics are reported per route, and how do requests carry adapter ids?
Data
How is the prompt laid out, and what KV memory budget does each replica have?
Architecture
Which levers apply in which order, and how are replicas routed and scaled?
Overview
The existing inference pages explain how one request runs: prefill vs decode, the KV cache, PagedAttention and continuous batching, the vLLM, SGLang and TensorRT-LLM runtimes, speculative decoding and multi-GPU parallelism. This cluster is about running a fleet for money and SLOs and adds four levers those pages only touch: weight quantization in depth (GPTQ, AWQ, SmoothQuant, FP8, INT4/INT8 and their accuracy trade-offs), prefix and prompt caching across requests (self-hosted block caches and provider prompt caching, cache keys, routing, TTL, hit-rate economics), multi-LoRA serving (many fine-tunes on one base), and latency and cost metrics with capacity planning (TTFT, TPOT/ITL, goodput, dollars per million tokens, GPU sizing). The hub shows which lever moves which number and ends with a triage chart.
Not re-taught here (see Related): prefill vs decode and KV mechanics, PagedAttention and continuous batching, vLLM and SGLang internals and choice, TensorRT-LLM engine builds, speculative decoding, tensor/expert parallelism and P/D disaggregation, constrained decoding, and the serving-stack choice playbook.
Four levers, four different bottlenecks
| Lever | Bottleneck it attacks | Moves | Costs or risks | Deep dive |
|---|---|---|---|---|
| Weight (and KV) quantization | Bytes per decode step, GPU memory | TPOT, tokens/s, GPU count, fits-on-GPU | Accuracy loss (small models, reasoning, long context first), kernel and hardware support | llm-weight-quantization-gptq-awq-smoothquant-fp8 |
| Prefix / prompt caching | Recomputing identical prompt prefixes | TTFT, prefill compute, input-token cost | Cache-busting prompt layouts, routing affinity, TTL economics, tenant isolation | prefix-prompt-caching-cache-keys-routing-ttl-economics |
| Multi-LoRA serving | One deployment per fine-tune | GPU count for many variants, utilisation | Adapter swap latency, kernel overhead, noisy neighbours | multi-lora-serving-slora-punica-adapter-hot-swap |
| Metrics and capacity planning | Guessing | The decision itself: how many GPUs, at what utilisation | Wrong metric (averages, TPOT hiding stalls), wrong workload model | llm-latency-cost-capacity-ttft-tpot-goodput-gpu-sizing |
The first three are mostly independent and stack: an FP8 base with FP8 KV, a prefix cache in front, and dozens of LoRA adapters on the same replica is a normal 2026 deployment. The fourth tells you whether you needed any of them.
Which lever first for a long shared prompt?
Prefer
Prefix caching (plus FP8 KV)
Skip prefill for the shared prefix and shrink KV bytes.
- TTFT dropped from 122 ms to 20 ms with a 5k-token shared prefix.
- FP8 KV cut TPOT from 14.4 ms to 8.9 ms.
- Freed memory let batch 128 fit at $0.17 per 1M.
Alternative
Weight quantization only
Fewer bytes per weight.
- FP8 weights halved TTFT from 244 ms to 122 ms but left KV untouched.
- W4A16 had worse TTFT (41 ms) than FP8 because prefill stays 16-bit.
- bf16 at batch 128 did not fit at all.
A production request meets the levers
Diagram 1 condensed: the request path through a fleet, with the failure path.
- 1
Gateway and routing
The router hashes the prompt prefix and picks a replica likely to hold it. - 2
Prefix cache lookup
Shared prefix blocks skip prefill. - 3
Adapter and quantized base
The request's LoRA runs on an FP8 or INT4 base. - 4
Batched decode
Batch size trades TPOT for cost per token. - 5
Out of KV memory
Preemption or queueing breaks the latency SLO.
How a production request meets the levers
Diagram 1: the request path through a production fleet, with the failure path.
Decisions
- 1
Step 1: request arrives at the gateway (tenant, model, adapter id, prompt)
- nextStep 2: router hashes the prompt prefix and picks a replica likely to hold it
- 2
Step 2: router hashes the prompt prefix and picks a replica likely to hold it
- nextStep 3: prefix blocks cached on that replica?
- ?
Step 3: prefix blocks cached on that replica?
- nextStep 4a: reuse cached KV, prefill only the new suffix
- nextStep 4b: full prefill, insert new blocks into the cache
- 4
Step 4a: reuse cached KV, prefill only the new suffix
- nextStep 5: LoRA adapter resolved, loaded into a GPU slot if cold
- 5
Step 4b: full prefill, insert new blocks into the cache
- nextStep 5: LoRA adapter resolved, loaded into a GPU slot if cold
- 6
Step 5: LoRA adapter resolved, loaded into a GPU slot if cold
- nextStep 6: decode in a continuous batch on quantized weights and FP8 KV
- 7
Step 6: decode in a continuous batch on quantized weights and FP8 KV
- nextStep 7: stream tokens, record TTFT, ITL, tokens and cost per request
- 8
Step 7: stream tokens, record TTFT, ITL, tokens and cost per request
- nextStep 8: SLOs met (goodput) at target utilisation?
- ?
Step 8: SLOs met (goodput) at target utilisation?
- nextStep 9: capacity plan holds, review weekly
- nextFailure path: p95 TTFT spikes, KV preemptions, adapter thrash
- 10
Step 9: capacity plan holds, review weekly
- 11
Failure path: p95 TTFT spikes, KV preemptions, adapter thrash
- nextTriage with the lever chart: caching for TTFT, quantization for memory and TPOT, slots for adapters, replicas for load
- 12
Triage with the lever chart: caching for TTFT, quantization for memory and TPOT, slots for adapters, replicas for load
- nextStep 2: router hashes the prompt prefix and picks a replica likely to hold it
Lesson map
LLM Inference in Production - Quantization, Prompt Caching, Multi-LoRA Serving & Capacity Planning
Hub: fleet-level serving levers beyond the existing runtime pages; which lever (quantization, prefix caching, multi-LoRA, batching/capacity) moves TTFT, TPOT, memory or $/1M; runnable lever simulation + symptom-to-lever picker; triage chart.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: request arrives at the gateway (tenant, model, adapter id, prompt)"] s2["Step 2: router hashes the prompt prefix and picks a replica likely to hold it"] s3["Step 3: prefix blocks cached on that replica?"] s4["Step 4a: reuse cached KV, prefill only the new suffix"] s5["Step 4b: full prefill, insert new blocks into the cache"] s6["Step 5: LoRA adapter resolved, loaded into a GPU slot if cold"] s7["Step 6: decode in a continuous batch on quantized weights and FP8 KV"] s8["Step 7: stream tokens, record TTFT, ITL, tokens and cost per request"] s9["Step 8: SLOs met (goodput) at target utilisation?"] s10["Step 9: capacity plan holds, review weekly"] f1["Failure path: p95 TTFT spikes, KV preemptions, adapter thrash"] f2["Triage with the lever chart: caching for TTFT, quantization for memory and TPOT, slots for adapters, replicas for load"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s3 -->|continues| s5 s4 -->|continues| s6 s5 -->|continues| s6 s6 -->|continues| s7 s7 -->|continues| s8 s8 -->|continues| s9 s9 -->|continues| s10 s9 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s2
Runnable: which lever moves which number
A deliberately simple roofline model of one GPU (example H100-class numbers) serving an 8B model with long prompts. It is not a benchmark; it shows the direction and rough size of each lever, including the failure row where a bigger batch does not fit.
"""Which production lever moves which number? A roofline-style model of one GPU.
SIMPLIFIED MODEL with labelled example numbers (not a benchmark):
* decode step time = bytes read per step / (HBM bandwidth * efficiency)
bytes = weights + KV cache of every sequence in the batch
* prefill time = 2 * params * uncached_prompt_tokens / (peak FLOPs * MFU)
* $ per 1M output tokens = GPU $/hour / output tokens per hour
Model shape: Llama-3-8B-like (8.03e9 params, 32 layers, 8 KV heads x 128 dims).
GPU: H100 SXM-class example numbers: 80 GB, 3.35 TB/s, ~989 dense BF16 TFLOPS, ~1979 dense FP8 TFLOPS.
"""
PARAMS, LAYERS, KV_HEADS, HEAD_DIM = 8.03e9, 32, 8, 128
HBM_GB, BW, BW_EFF = 80, 3.35e12, 0.7
PEAK = {"bf16": 989e12, "fp8": 1979e12}
MFU = 0.4
GPU_PER_HOUR = 3.00 # example on-demand price, USD
PROMPT, SHARED_PREFIX, OUTPUT, BATCH = 6000, 5000, 300, 32
def scenario(name, w_bytes=2, kv_bytes=2, prefix_cache=False, compute="bf16", batch=BATCH):
weights = PARAMS * w_bytes
kv_per_token = 2 * LAYERS * KV_HEADS * HEAD_DIM * kv_bytes # K and V
avg_ctx = PROMPT + OUTPUT / 2
kv_read = batch * avg_ctx * kv_per_token # attention reads every sequence's KV
# With a prefix cache, shared blocks are STORED once (memory), though still READ per sequence.
kv_mem = kv_read if not prefix_cache else (SHARED_PREFIX + batch * (avg_ctx - SHARED_PREFIX)) * kv_per_token
fits = weights + kv_mem <= HBM_GB * 1e9 * 0.9
step = (weights + kv_read) / (BW * BW_EFF) # seconds per decode step
tpot_ms = step * 1e3
tok_s = batch / step
uncached = PROMPT - (SHARED_PREFIX if prefix_cache else 0)
ttft_ms = 2 * PARAMS * uncached / (PEAK[compute] * MFU) * 1e3
usd = GPU_PER_HOUR / (tok_s * 3600) * 1e6
print(f"{name:<34}{weights/1e9:>6.1f}{kv_mem/1e9:>7.1f}{'yes' if fits else 'NO':>5}"
f"{ttft_ms:>8.0f}{tpot_ms:>7.1f}{tok_s:>8.0f}{usd:>8.2f}")
print(f"{'scenario':<34}{'W GB':>6}{'KV GB':>7}{'fit':>5}{'TTFT':>8}{'TPOT':>7}{'tok/s':>8}{'$/1M':>8}")
scenario("baseline bf16 weights + bf16 KV")
scenario("+ FP8 weights (W8A8)", w_bytes=1, compute="fp8")
scenario("+ FP8 KV cache", w_bytes=1, kv_bytes=1, compute="fp8")
scenario("+ prefix cache (5k shared tokens)", w_bytes=1, kv_bytes=1, compute="fp8", prefix_cache=True)
scenario(" + batch 128 in the freed memory", w_bytes=1, kv_bytes=1, compute="fp8", prefix_cache=True, batch=128)
scenario("baseline bf16 at batch 128", batch=128)
scenario("INT4 weights only (W4A16) + cache", w_bytes=0.5, kv_bytes=1, prefix_cache=True)
print("\nKV GB = KV memory held; TTFT = ms of prefill compute only (no queueing); TPOT = ms per output token.")
print("Long prompts make the KV cache, not the weights, the dominant bytes per decode step, so FP8 KV and")
print("prefix caching matter as much as weight quantization. Bigger batches buy throughput ($/1M) with TPOT;")
print("the failure path is the bf16 batch-128 row: it does not fit, so the server would preempt or queue.")
print("Note: W4A16 computes prefill in 16-bit, so its TTFT is worse than FP8 even though its weights are smaller.")Output:
scenario W GB KV GB fit TTFT TPOT tok/s $/1M
baseline bf16 weights + bf16 KV 16.1 25.8 yes 244 17.8 1793 0.46
+ FP8 weights (W8A8) 8.0 25.8 yes 122 14.4 2218 0.38
+ FP8 KV cache 8.0 12.9 yes 122 8.9 3586 0.23
+ prefix cache (5k shared tokens) 8.0 2.7 yes 20 8.9 3586 0.23
+ batch 128 in the freed memory 8.0 10.0 yes 20 25.4 5035 0.17
baseline bf16 at batch 128 16.1 103.2 NO 244 50.8 2517 0.33
INT4 weights only (W4A16) + cache 4.0 2.7 yes 41 7.2 4437 0.19
KV GB = KV memory held; TTFT = ms of prefill compute only (no queueing); TPOT = ms per output token.
Long prompts make the KV cache, not the weights, the dominant bytes per decode step, so FP8 KV and
prefix caching matter as much as weight quantization. Bigger batches buy throughput ($/1M) with TPOT;
the failure path is the bf16 batch-128 row: it does not fit, so the server would preempt or queue.
Note: W4A16 computes prefill in 16-bit, so its TTFT is worse than FP8 even though its weights are smaller.What the table teaches:
- With long prompts, the KV cache, not the weights, becomes the dominant bytes per decode step, so FP8 KV matters as much as FP8 weights.
- Prefix caching collapses TTFT when prompts share a long prefix, and storing shared blocks once frees memory for bigger batches.
- W4A16 (4-bit weights, 16-bit compute) wins decode speed and memory but not prefill, which is compute-bound.
- Bigger batches lower dollars per token and raise TPOT; capacity planning is choosing where on that curve your SLO lets you sit.
Runnable: symptom-to-lever triage
// Symptom -> lever triage for an LLM serving fleet (mirrors the hub decision chart).
// Input is what your dashboards show; output is the ordered list of levers to try.
type Fleet = {
ttftP95Ms: number; ttftSloMs: number;
tpotP95Ms: number; tpotSloMs: number;
sharedPrefixFraction: number; // share of prompt tokens identical across requests (from logs)
kvCacheUsage: number; // 0..1, from the runtime's KV usage gauge
preemptionsPerMin: number; // sequences evicted/recomputed for lack of KV memory
fineTunedVariants: number; // distinct fine-tunes you must serve
qualityHeadroom: boolean; // does the eval suite tolerate a small accuracy loss?
};
function triage(f: Fleet): string[] {
const out: string[] = [];
if (f.ttftP95Ms > f.ttftSloMs) {
if (f.sharedPrefixFraction > 0.3) out.push("TTFT: enable prefix caching + prefix-aware routing (skip recomputing shared tokens)");
out.push("TTFT: check queueing first (utilisation too high?) then chunked prefill / P-D disaggregation (see inference parallelism page)");
}
if (f.kvCacheUsage > 0.9 || f.preemptionsPerMin > 0) {
out.push("Memory: FP8 KV cache, then FP8/INT4 weights, to free HBM for KV (bigger batch, fewer preemptions)");
}
if (f.tpotP95Ms > f.tpotSloMs) {
out.push(f.qualityHeadroom
? "TPOT: weight quantization cuts bytes per decode step (memory-bound); verify on evals"
: "TPOT: lower max batch / add replicas; speculative decoding if acceptance is high (see its page)");
}
if (f.fineTunedVariants > 3) {
out.push(`Fleet: ${f.fineTunedVariants} fine-tunes -> serve LoRA adapters on one base (multi-LoRA) instead of ${f.fineTunedVariants} deployments`);
}
if (out.length === 0) out.push("Within SLO: capacity-plan headroom and cost per 1M tokens instead of tuning");
return out;
}
const fleets: Array<[string, Fleet]> = [
["RAG chatbot, long system prompt", { ttftP95Ms: 2100, ttftSloMs: 800, tpotP95Ms: 30, tpotSloMs: 50, sharedPrefixFraction: 0.7, kvCacheUsage: 0.95, preemptionsPerMin: 12, fineTunedVariants: 1, qualityHeadroom: true }],
["B2B SaaS, 40 customer fine-tunes", { ttftP95Ms: 400, ttftSloMs: 800, tpotP95Ms: 70, tpotSloMs: 50, sharedPrefixFraction: 0.1, kvCacheUsage: 0.6, preemptionsPerMin: 0, fineTunedVariants: 40, qualityHeadroom: false }],
["healthy internal tool", { ttftP95Ms: 300, ttftSloMs: 800, tpotP95Ms: 25, tpotSloMs: 50, sharedPrefixFraction: 0.2, kvCacheUsage: 0.5, preemptionsPerMin: 0, fineTunedVariants: 1, qualityHeadroom: true }],
];
for (const [name, f] of fleets) {
console.log(`\n# ${name}`);
triage(f).forEach((l, i) => console.log(` ${i + 1}. ${l}`));
}Output:
# RAG chatbot, long system prompt
1. TTFT: enable prefix caching + prefix-aware routing (skip recomputing shared tokens)
2. TTFT: check queueing first (utilisation too high?) then chunked prefill / P-D disaggregation (see inference parallelism page)
3. Memory: FP8 KV cache, then FP8/INT4 weights, to free HBM for KV (bigger batch, fewer preemptions)
# B2B SaaS, 40 customer fine-tunes
1. TPOT: lower max batch / add replicas; speculative decoding if acceptance is high (see its page)
2. Fleet: 40 fine-tunes -> serve LoRA adapters on one base (multi-LoRA) instead of 40 deployments
# healthy internal tool
1. Within SLO: capacity-plan headroom and cost per 1M tokens instead of tuningExpected# RAG chatbot, long system prompt 1. TTFT: enable prefix caching + prefix-aware routing (skip recomputing shared tokens) 2. TTFT: check queueing first (utilisation too high?) then chunked prefill / P-D disaggregation (see inference parallelism page) 3. Memory: FP8 KV cache, then FP8/INT4 weights, to free HBM for KV (bigger batch, fewer preemptions) # B2B SaaS, 40 customer fine-tunes 1. TPOT: lower max batch / add replicas; speculative decoding if acceptance is high (see its page) 2. Fleet: 40 fine-tunes -> serve LoRA adapters on one base (multi-LoRA) instead of 40 deployments # healthy internal tool 1. Within SLO: capacity-plan headroom and cost per 1M tokens instead of tuning
Press Run. Snippets must be self-contained — no network, files, or native modules.
Decision chart: which lever first?
Decisions
- 1
D1
- nextOptimise cost: quantize, raise batch within TPOT SLO, consolidate fine-tunes
- nextStep 2: is TTFT the broken metric?
- 2
Optimise cost: quantize, raise batch within TPOT SLO, consolidate fine-tunes
- nextStep 6: many fine-tuned variants?
- ?
Step 2: is TTFT the broken metric?
- nextStep 3: do prompts share long prefixes?
- nextStep 4: KV memory full or preemptions?
- ?
Step 3: do prompts share long prefixes?
- nextPrefix caching plus prefix-aware routing
- nextReduce queueing (replicas, admission), chunked prefill or P/D disaggregation
- 5
Prefix caching plus prefix-aware routing
- 6
Reduce queueing (replicas, admission), chunked prefill or P/D disaggregation
- ?
Step 4: KV memory full or preemptions?
- nextFP8 KV cache, then FP8 or INT4 weights
- nextStep 5: TPOT high with headroom in quality evals?
- 8
FP8 KV cache, then FP8 or INT4 weights
- ?
Step 5: TPOT high with headroom in quality evals?
- nextWeight quantization (memory-bound decode)
- nextLower max batch or add replicas, consider speculative decoding
- 10
Weight quantization (memory-bound decode)
- 11
Lower max batch or add replicas, consider speculative decoding
- ?
Step 6: many fine-tuned variants?
- nextMulti-LoRA serving on a shared base
- nextKeep one merged model per deployment
- 13
Multi-LoRA serving on a shared base
- 14
Keep one merged model per deployment
What happens if you choose otherwise
- Quantize first when TTFT is the problem: weight quantization barely changes queueing or prefill for shared prompts; prefix caching would have.
- Turn on prefix caching with a timestamp at the top of the prompt: hit rate stays near zero and nothing changes.
- Serve 40 customer fine-tunes as 40 deployments: dozens of mostly idle GPUs; multi-LoRA serves them on a few.
- Plan capacity from average throughput: the fleet looks 60 percent utilised while p95 TTFT already breaks the SLO at peak.
Pitfalls
- Benchmarking with synthetic prompts that share nothing (or everything) and then generalising cache results to production.
- Comparing quantized vs bf16 on perplexity only; run task evals, long-context and reasoning sets.
- Mixing up per-request TPOT with fleet tokens per second.
- Forgetting that every lever changes the others: quantization frees memory that caching and batching then consume.
The cluster map
- LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offs: RTN and scale granularity, outliers and WxAy notation, GPTQ vs AWQ, SmoothQuant W8A8, FP8 E4M3, calibration and when quantization helps TTFT vs only TPOT.
- Prefix & Prompt Caching Across Requests - Cache Keys, Routing, TTL & Hit-Rate Economics: KV block hashing vs radix trees, what goes into cache keys, prefix-aware routing across replicas, provider prompt caching and hit-rate economics.
- Multi-LoRA Serving - S-LoRA/Punica Batching, Adapter Hot-Swap & Memory: merged deployments vs in-place merging vs multi-LoRA batching, SGMV/BGMV kernels, S-LoRA paging, adapter tiers, hot-swap and affinity routing.
- LLM Latency, Cost & Capacity Planning - TTFT, TPOT/ITL, Goodput, $/1M Tokens & GPU Sizing: TTFT, TPOT, ITL and goodput, a worked 70B FP8 sizing, the Erlang C TTFT hockey stick, and autoscaling signals compared.
Interview Q&A
You inherit an LLM fleet that is over budget. Where do you start?
Answer
Measure first: per-request TTFT, TPOT/ITL, tokens, cache hit rate and cost per million tokens, by route. Then pick levers by bottleneck: quantization for memory and decode bandwidth, prefix caching where prompts share prefixes, multi-LoRA to consolidate fine-tunes, and a utilisation target derived from p95 SLOs.
Which lever helps TTFT and which helps TPOT?
Answer
Prefix caching, queueing control and chunked prefill or P/D disaggregation help TTFT; weight and KV quantization, batch size limits and speculative decoding help TPOT. Quantization helps TTFT only when it also speeds compute (FP8 W8A8), not with weight-only W4A16.
Why can FP8 KV matter as much as FP8 weights?
Answer
With long contexts and large batches, KV bytes read per decode step exceed weight bytes, so halving KV bytes saves more bandwidth and memory than halving weights.
When is multi-LoRA serving worth it?
Answer
When you have several fine-tunes of the same base with uneven traffic. One base replica serves all of them with small per-adapter overhead instead of one mostly idle deployment per fine-tune.
You inherit an over-budget fleet. Where do you start?
Answer
Measure TTFT, TPOT/ITL, tokens, cache hit rate and cost per million tokens by route, then pick levers by bottleneck.
In the lever simulation, why did bf16 at batch 128 fail?
Answer
Its KV needed 103.2 GB, so it did not fit; the server would preempt or queue.
Why do the first three levers stack?
Answer
They are mostly independent: an FP8 base with FP8 KV, a prefix cache and many LoRA adapters on one replica is a normal 2026 deployment.
Why plan capacity from p95 SLOs rather than average throughput?
Answer
A fleet can look 60 percent utilised on average while p95 TTFT already breaks the SLO at peak.
Check yourself
For one route you serve, write its TTFT, TPOT and cost per 1M tokens, the share of prompt tokens that are a shared prefix, and the number of fine-tunes. Pick the first lever with the triage chart.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: LLM Inference Runtime — Prefill, KV, Batching & Parallelism, LLM Inference Engines & Distributed Training — SGLang, TensorRT-LLM, ONNX & PyTorch Parallelism, Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes, vLLM vs SGLang — Runtime Choice (comparative), Prefill vs Decode & KV Cache Mechanics, Speculative Decoding, Inference Parallelism — Tensor Parallel, Expert Parallel & Prefill/Decode Disaggregation, LLM Post-Training - From Base Model to Assistant: SFT, Preference Tuning, PEFT, Distillation & Evals, LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter Merging.