LLM Latency, Cost & Capacity Planning - TTFT, TPOT/ITL, Goodput, $/1M Tokens & GPU Sizing
TTFT vs TPOT vs ITL, goodput, percentiles from histograms; worked 70B FP8 fleet sizing (KV memory, prefill share, Little's law, headroom, $/1M); Erlang C TTFT hockey stick; autoscaling signals compared; runnable capacity plan + TTFT calculator.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is goodput?
Answer
Requests per second that meet all latency SLOs.
L2
TPOT vs ITL?
Answer
TPOT is a per-request average; ITL is the distribution of gaps between tokens.
L3
How do you bound concurrency?
Answer
Memory left after weights divided by KV bytes per token and sequence length.
L4
What does Little's law give?
Answer
Sequences decoding = arrival rate x output tokens x effective TPOT.
L5
Why not run at 95 percent utilisation?
Answer
Queueing delay grows sharply near 100 percent, so p95 TTFT explodes.
L6
What should you autoscale on?
Answer
Queue depth and KV-cache usage as leading signals, p95 TTFT as a guardrail.
L7
How do you quote cost per 1M tokens?
Answer
GPU dollars per hour over tokens per hour at the planned utilisation, input and output separately.
Failure modes
Plan from average tokens per second
The fleet fails p95 TTFT at every peak.
Autoscale on GPU utilisation
Continuous batching keeps it high at any load, so scaling triggers late.
Uniform short benchmark prompts
Long RAG prompts shift the bottleneck to prefill and halve capacity.
Misconceptions
TPOT shows stalls.
It averages them away; ITL p99 and max show them.
Throughput is the metric to maximise.
Throughput can rise while SLO compliance collapses.
Decode bandwidth always runs out first.
With long prompts and short answers, prefill compute saturates first.
Interviewer traps
Forgetting N+1 capacity.
Add capacity for replica failures and rolling deploys.
Mixing batch and interactive traffic without priorities.
Batch jobs eat the headroom interactive SLOs need.
Design scenario
Same prompt for every reader.
Requirements
Meet p95 TTFT and TPOT SLOs at peak, survive a replica failure, and quote cost per 1M tokens.
Failure assumptions
- Traffic is bursty.
- About 5 percent of streams stall mid-answer.
- Cold starts take minutes.
Constraints
- 4-GPU replicas.
- An example price of $3 per GPU-hour.
Prompt
Size a fleet for a Llama-3-70B-like model in FP8 at 20 req/s with 2,000-token prompts and 400-token answers.
API
Which SLOs, percentiles and goodput definitions does the API commit to?
Data
Which workload model (arrivals, lengths, cache hit rate) and per-replica measurements feed the plan?
Architecture
How many replicas, at what utilisation target, scaled on which signals?
Overview
Capacity planning for LLM serving answers one question: how many GPUs of which kind, at what utilisation, so that a stated share of requests meets latency SLOs at a known cost per token. You need four things: the right metrics (TTFT, TPOT vs ITL, tokens per second at fleet and request level, goodput), a workload model (arrival rate, prompt and output length distributions, cache hit rate), a performance model of one replica (prefill compute, decode bandwidth, KV memory), and queueing to turn utilisation into tail latency. This page builds each, works through a full 70B sizing example, and shows the failure modes: averages hiding stalls, throughput without SLOs, and autoscaling on the wrong signal. TTFT, TPOT and the KV-per-token formula are introduced on the prefill-vs-decode page, and the serving playbook has a basic cost-per-1M snippet. This page goes deeper into tails, goodput, queueing and sizing.
Metrics: what each one catches
| Metric | Definition | Catches | Misses |
|---|---|---|---|
| TTFT | Request arrival to first token | Queueing, prefill cost, cache misses | Stalls later in the stream |
| TPOT | (end-to-end latency minus TTFT) / (output tokens - 1), averaged per request | Steady decode speed | Individual stalls (it is an average) |
| ITL (inter-token latency) | Gap between consecutive tokens, as a distribution | Stalls from prefill interference, preemption, swaps | Nothing per token; it is the raw signal |
| E2E latency | Arrival to last token | What non-streaming clients feel | Where the time went |
| Output tokens/s (fleet) | All generated tokens per second | Throughput, cost basis | Whether any user was served well |
| Requests/s at SLO (goodput) | Requests per second meeting all SLOs (for example p95 TTFT and TPOT) | The useful capacity | - |
| Cost per 1M tokens | GPU dollars per hour / tokens per hour | Unit economics | Quality and SLO violations |
Faster decode kernels or less prefill work?
Prefer
Cut prefill work (prefix caching, chunked prefill, P/D)
Prefill compute runs out first with long prompts.
- At 40 req/s, a 60% prefix-cache hit took p95 TTFT from 607 ms to 38 ms.
- Utilisation dropped from 88% to 35%.
- Capacity levers compound.
Alternative
Faster decode kernels
Lower decode step time.
- Per-replica max stayed about 8 req/s, bound by prefill share.
- At 9 req/s prefill share was 80% and TPOT 113.4 ms, over SLO.
- Decode speed does not fix a prefill-bound replica.
The capacity planning loop
Diagram 1 condensed: SLOs to a fleet size, with the failure path.
- 1
Fix SLOs and peak workload
p95 TTFT and TPOT, arrival rate, prompt and output lengths. - 2
Model one replica
KV memory, prefill share, decode step time. - 3
Solve with Little's law
Max req/s per replica within SLO. - 4
Add headroom and N+1
Utilisation target from the p95 TTFT curve. - 5
Load test breaks early
Check the workload model (longer prompts, bursts, low cache hits) and engine settings.
Goodput (popularised by DistServe) is the metric to optimise. A configuration can double raw throughput by running huge batches while half the requests miss the TPOT SLO; goodput falls even though throughput rose.
Report percentiles from histograms, not averages, and report input and output tokens separately: prefill and decode cost different resources.
"""LLM latency metrics from token timestamps, goodput, and a worked GPU sizing example.
Part 1 turns raw streaming timestamps into TTFT, TPOT, ITL and goodput.
Part 2 sizes a fleet for a target load with a SIMPLIFIED roofline model and labelled
example hardware numbers (H100 SXM-class: 80 GB, 3.35 TB/s, ~1979 dense FP8 TFLOPS).
"""
import math, random
random.seed(5)
# ---------- Part 1: metrics from timestamps ----------
def pct(xs, p):
xs = sorted(xs); k = (len(xs) - 1) * p / 100
f, c = math.floor(k), math.ceil(k)
return xs[f] + (xs[c] - xs[f]) * (k - f)
reqs = []
for i in range(400):
sent = 0.0
queue = random.expovariate(1 / 0.15) if random.random() < 0.9 else random.uniform(1.0, 3.0) # 10% hit a queue spike
ttft = queue + 0.12 + random.random() * 0.05
n = random.randint(50, 400)
gaps = [max(0.005, random.gauss(0.03, 0.006)) for _ in range(n - 1)]
if random.random() < 0.05: # a prefill of someone else's long prompt stalls this stream
gaps[random.randrange(n - 1)] += 0.8
times = [sent + ttft]
for g in gaps: times.append(times[-1] + g)
reqs.append(times)
TTFT_SLO, TPOT_SLO = 1.0, 0.05
ttfts = [t[0] for t in reqs]
tpots = [(t[-1] - t[0]) / (len(t) - 1) for t in reqs] # mean time per output token after the first
itls = [b - a for t in reqs for a, b in zip(t, t[1:])] # every inter-token gap
good = sum(1 for a, b in zip(ttfts, tpots) if a <= TTFT_SLO and b <= TPOT_SLO)
print(f"TTFT p50={pct(ttfts,50)*1e3:.0f}ms p95={pct(ttfts,95)*1e3:.0f}ms p99={pct(ttfts,99)*1e3:.0f}ms")
print(f"TPOT p50={pct(tpots,50)*1e3:.1f}ms p95={pct(tpots,95)*1e3:.1f}ms ITL p99={pct(itls,99)*1e3:.1f}ms max={max(itls)*1e3:.0f}ms")
print(f"goodput: {good}/{len(reqs)} requests met BOTH SLOs ({good/len(reqs):.0%}); throughput counts all {len(reqs)}")
print("TPOT averages away stalls; ITL max/p99.9 shows the 0.8 s freeze a user actually sees.\n")
# ---------- Part 2: size a fleet ----------
P = 70e9; LAYERS, KV_HEADS, HEAD_DIM = 80, 8, 128 # Llama-3-70B-like
W_BYTES, KV_BYTES = 1, 1 # FP8 weights and FP8 KV cache
TP, HBM, BW, PEAK, MFU, BW_EFF = 4, 80e9, 3.35e12, 1979e12, 0.4, 0.7
RPS, IN_TOK, OUT_TOK = 20, 2000, 400
GPU_HOUR, TARGET_UTIL = 3.00, 0.7 # example price, utilisation headroom
kv_tok = 2 * LAYERS * KV_HEADS * HEAD_DIM * KV_BYTES # bytes of KV per token
weights = P * W_BYTES
ctx = IN_TOK + OUT_TOK / 2
max_batch_mem = int((TP * HBM * 0.9 - weights) / (ctx * kv_tok))
step = lambda b: (weights + b * ctx * kv_tok) / (TP * BW * BW_EFF) # decode step seconds
prefill_s = 2 * P * IN_TOK / (TP * PEAK * MFU) # replica-seconds per prompt
print(f"replica = {TP} GPUs; KV/token={kv_tok/1024:.0f} KiB; memory allows {max_batch_mem} concurrent sequences")
TPOT_SLO_S, HEADROOM = 0.05, 0.8
def replica_state(lam):
"""Per-replica steady state at lam req/s. Prefill steals time from decode (u_p), so effective
TPOT = step(b) / (1 - u_p); Little's law: sequences decoding b = lam * OUT_TOK * TPOT_eff."""
u_p = lam * prefill_s
if u_p >= 1:
return None
b = 1.0
for _ in range(200):
b_new = lam * OUT_TOK * step(b) / (1 - u_p)
if b_new > max_batch_mem:
return None # KV memory exhausted -> preemption / queue grows
if abs(b_new - b) < 1e-6:
break
b = b_new
else:
return None
return u_p, b, step(b) / (1 - u_p)
print(f"{'req/s per replica':>17}{'prefill share':>14}{'seqs decoding':>14}{'eff. TPOT':>11}")
best = 0.0
for lam in [2, 4, 6, 7, 8, 9, 10]:
st = replica_state(lam)
if st is None:
print(f"{lam:>17} unstable: no steady state (queue grows)"); continue
u_p, b, tpot = st
ok = tpot <= TPOT_SLO_S
if ok: best = max(best, lam)
print(f"{lam:>17}{u_p:>14.0%}{b:>14.0f}{tpot*1e3:>9.1f}ms{'' if ok else ' > SLO'}")
replicas = math.ceil(RPS / (best * HEADROOM))
print(f"max load per replica within TPOT SLO ~{best} req/s; with {HEADROOM:.0%} headroom -> {replicas} replicas = {replicas*TP} GPUs")
cost_h = replicas * TP * GPU_HOUR
out_per_h, all_per_h = RPS * OUT_TOK * 3600, RPS * (IN_TOK + OUT_TOK) * 3600
print(f"cost ${cost_h:.0f}/h -> ${cost_h/out_per_h*1e6:.2f} per 1M output tokens, ${cost_h/all_per_h*1e6:.2f} per 1M total tokens")
print("Prefill compute, not decode bandwidth, is what runs out first here (long prompts, short answers):")
print("as prefill share approaches 100% the effective TPOT explodes. Prefix caching, shorter prompts,")
print("chunked prefill or P/D disaggregation buy more capacity than faster decode kernels would.")Output:
TTFT p50=257ms p95=2324ms p99=2995ms
TPOT p50=30.0ms p95=32.4ms ITL p99=44.1ms max=840ms
goodput: 349/400 requests met BOTH SLOs (87%); throughput counts all 400
TPOT averages away stalls; ITL max/p99.9 shows the 0.8 s freeze a user actually sees.
replica = 4 GPUs; KV/token=160 KiB; memory allows 604 concurrent sequences
req/s per replica prefill share seqs decoding eff. TPOT
2 18% 8 9.4ms
4 35% 20 12.8ms
6 53% 47 19.8ms
7 62% 76 27.3ms
8 71% 141 44.0ms
9 80% 408 113.4ms > SLO
10 unstable: no steady state (queue grows)
max load per replica within TPOT SLO ~8 req/s; with 80% headroom -> 4 replicas = 16 GPUs
cost $48/h -> $1.67 per 1M output tokens, $0.28 per 1M total tokens
Prefill compute, not decode bandwidth, is what runs out first here (long prompts, short answers):
as prefill share approaches 100% the effective TPOT explodes. Prefix caching, shorter prompts,
chunked prefill or P/D disaggregation buy more capacity than faster decode kernels would.The worked example, explained
Part 1 computes metrics from per-token timestamps of 400 synthetic requests, where about 10 percent hit a queue spike and about 5 percent of streams freeze for 0.8 s mid-answer. TPOT barely notices the freezes; ITL max does; goodput counts only requests that met both SLOs.
Part 2 sizes a fleet for a Llama-3-70B-like model in FP8 (weights and KV) on 4-GPU replicas at 20 req/s, 2,000-token prompts and 400-token answers:
- KV per token = 2 x 80 layers x 8 KV heads x 128 dims x 1 byte = 160 KiB, so memory left after weights holds about 600 concurrent sequences.
- Prefill costs 2 x params x prompt tokens of FLOPs at an assumed 40 percent MFU, which gives the share of each second the replica spends in prefill at a given arrival rate.
- Decode step time = (weight bytes + KV bytes of the batch) / effective bandwidth. Prefill steals time from decode, so effective TPOT = step / (1 - prefill share).
- Little's law links them: sequences decoding = arrival rate x output tokens x effective TPOT. Solve for the steady state at each arrival rate.
- The highest rate within the TPOT SLO is about 8 req/s per replica. With 80 percent headroom, 20 req/s needs 4 replicas, which is 16 GPUs, and at an example $3 per GPU-hour that comes to about $1.67 per 1M output tokens.
The key insight is in the last line: with long prompts and short answers, prefill compute saturates first. Prefix caching, chunked prefill or P/D disaggregation buy more capacity here than faster decode kernels.
Queueing: why the utilisation target is not 95 percent
Requests arrive randomly. As utilisation approaches 1, waiting time grows without bound (M/M/c behaviour, Erlang C). TTFT is mostly queue wait plus prefill, so the p95 TTFT SLO decides the utilisation you can run at.
// TTFT calculator: TTFT = admission queue wait + prefill compute (+ network, ignored here).
// Queue model: M/M/c (Erlang C) over c "prefill lanes" - an approximation of a fleet where each
// replica admits one prefill at a time. Prefill time from FLOPs: 2 * params * prompt_tokens / (peak * MFU).
// Example numbers: 70B model, 4-GPU replica, ~1979 dense FP8 TFLOPS per GPU, MFU 40%.
function erlangC(c: number, a: number): number {
// probability an arrival waits; a = offered load in Erlangs (lambda / mu), requires a < c
let sum = 0, term = 1;
for (let k = 0; k < c; k++) { if (k > 0) term *= a / k; sum += term; }
const top = term * (a / c) * (c / (c - a)); // a^c / c! * c/(c-a)
return top / (sum + top);
}
function ttft(promptTokens: number, rps: number, replicas: number, cachedFraction = 0) {
const params = 70e9, peak = 4 * 1979e12, mfu = 0.4;
const prefill = (2 * params * promptTokens * (1 - cachedFraction)) / (peak * mfu); // seconds
const mu = 1 / prefill, a = rps / mu;
if (a >= replicas) return { util: a / replicas, p50: Infinity, p95: Infinity };
const pw = erlangC(replicas, a);
const rate = replicas * mu - rps; // P(W > t) = pw * exp(-rate * t)
const q = (p: number) => (pw <= 1 - p ? 0 : Math.log(pw / (1 - p)) / rate);
return { util: a / replicas, p50: q(0.5) + prefill, p95: q(0.95) + prefill };
}
const ms = (s: number) => (Number.isFinite(s) ? `${(s * 1e3).toFixed(0)} ms` : "unbounded");
console.log("prompt 2000 tok, 4 replicas: TTFT vs load (prefill lanes only)");
console.log("req/s prefill-util TTFT p50 TTFT p95");
for (const rps of [10, 20, 30, 36, 40, 44, 46]) {
const r = ttft(2000, rps, 4);
console.log(`${String(rps).padEnd(7)}${(100 * r.util).toFixed(0).padStart(6)}% ${ms(r.p50).padEnd(13)}${ms(r.p95)}`);
}
console.log("\nsame 40 req/s, with 60% of prompt tokens served from a prefix cache:");
const c = ttft(2000, 40, 4, 0.6);
console.log(` util ${(100 * c.util).toFixed(0)}% TTFT p50 ${ms(c.p50)} p95 ${ms(c.p95)}`);
console.log("\nThe hockey stick: p95 TTFT grows gently until roughly 65-80% utilisation, then queueing dominates and");
console.log("explodes as utilisation -> 100%. Capacity plans pick a utilisation target from the p95 SLO, not from");
console.log("average throughput. Failure path: autoscaling on average GPU utilisation reacts after p95 already broke.");Output:
prompt 2000 tok, 4 replicas: TTFT vs load (prefill lanes only)
req/s prefill-util TTFT p50 TTFT p95
10 22% 88 ms 88 ms
20 44% 88 ms 124 ms
30 66% 88 ms 220 ms
36 80% 106 ms 356 ms
40 88% 167 ms 607 ms
44 97% 600 ms 2466 ms
46 102% unbounded unbounded
same 40 req/s, with 60% of prompt tokens served from a prefix cache:
util 35% TTFT p50 35 ms p95 38 ms
The hockey stick: p95 TTFT grows gently until roughly 65-80% utilisation, then queueing dominates and
explodes as utilisation -> 100%. Capacity plans pick a utilisation target from the p95 SLO, not from
average throughput. Failure path: autoscaling on average GPU utilisation reacts after p95 already broke.Expectedprompt 2000 tok, 4 replicas: TTFT vs load (prefill lanes only) req/s prefill-util TTFT p50 TTFT p95 10 22% 88 ms 88 ms 20 44% 88 ms 124 ms 30 66% 88 ms 220 ms 36 80% 106 ms 356 ms 40 88% 167 ms 607 ms 44 97% 600 ms 2466 ms 46 102% unbounded unbounded same 40 req/s, with 60% of prompt tokens served from a prefix cache: util 35% TTFT p50 35 ms p95 38 ms The hockey stick: p95 TTFT grows gently until roughly 65-80% utilisation, then queueing dominates and explodes as utilisation -> 100%. Capacity plans pick a utilisation target from the p95 SLO, not from average throughput. Failure path: autoscaling on average GPU utilisation reacts after p95 already broke.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The same 40 req/s that sits on the steep part of the curve becomes trivial with a 60 percent prefix-cache hit rate, because caching cut the prefill work per request. Capacity levers compound.
Diagram 1: the capacity planning loop
Decisions
- 1
Step 1: define SLOs - p95 TTFT, p95 TPOT or p99 ITL, availability - and the goodput target
- nextStep 2: model the workload - arrival rate at peak, prompt and output length distributions, cache hit rate
- 2
Step 2: model the workload - arrival rate at peak, prompt and output length distributions, cache hit rate
- nextStep 3: model one replica - memory for weights and KV, prefill FLOPs, decode bandwidth
- 3
Step 3: model one replica - memory for weights and KV, prefill FLOPs, decode bandwidth
- nextStep 4: find max req/s per replica that meets SLOs (steady state plus queueing)
- 4
Step 4: find max req/s per replica that meets SLOs (steady state plus queueing)
- nextStep 5: replicas = peak load / (max per replica x headroom), add N+1 for failures
- 5
Step 5: replicas = peak load / (max per replica x headroom), add N+1 for failures
- nextStep 6: load test with a production-shaped trace, compare goodput with the model
- 6
Step 6: load test with a production-shaped trace, compare goodput with the model
- nextStep 7: model and test agree within about 20 percent?
- ?
Step 7: model and test agree within about 20 percent?
- nextStep 8: deploy, autoscale on queue depth and TTFT, report cost per 1M tokens
- nextFailure path: test shows p95 TTFT breaking at half the predicted load
- 8
Step 8: deploy, autoscale on queue depth and TTFT, report cost per 1M tokens
- 9
Failure path: test shows p95 TTFT breaking at half the predicted load
- nextCheck the workload model (real prompts longer, bursty arrivals, low cache hits) and engine settings (max batch, chunked prefill)
- 10
Check the workload model (real prompts longer, bursty arrivals, low cache hits) and engine settings (max batch, chunked prefill)
- nextStep 2: model the workload - arrival rate at peak, prompt and output length distributions, cache hit rate
Lesson map
LLM Latency, Cost & Capacity Planning - TTFT, TPOT/ITL, Goodput, $/1M Tokens & GPU Sizing
TTFT vs TPOT vs ITL, goodput, percentiles from histograms; worked 70B FP8 fleet sizing (KV memory, prefill share, Little's law, headroom, $/1M); Erlang C TTFT hockey stick; autoscaling signals compared; runnable capacity plan + TTFT calculator.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: define SLOs - p95 TTFT, p95 TPOT or p99 ITL, availability - and the goodput target"] s2["Step 2: model the workload - arrival rate at peak, prompt and output length distributions, cache hit rate"] s3["Step 3: model one replica - memory for weights and KV, prefill FLOPs, decode bandwidth"] s4["Step 4: find max req/s per replica that meets SLOs (steady state plus queueing)"] s5["Step 5: replicas = peak load / (max per replica x headroom), add N+1 for failures"] s6["Step 6: load test with a production-shaped trace, compare goodput with the model"] s7["Step 7: model and test agree within about 20 percent?"] s8["Step 8: deploy, autoscale on queue depth and TTFT, report cost per 1M tokens"] f1["Failure path: test shows p95 TTFT breaking at half the predicted load"] f2["Check the workload model (real prompts longer, bursty arrivals, low cache hits) and engine settings (max batch, chunked prefill)"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s4 -->|continues| s5 s5 -->|continues| s6 s6 -->|continues| s7 s7 -->|continues| s8 s7 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s2
Sizing heuristics that hold up
- Memory first: weights + KV for the target concurrency + activations + about 10 percent overhead must fit, or nothing else matters.
- Decode ceiling: at small batch, tokens/s per sequence is at most bandwidth / bytes read per step.
- Prefill ceiling: prompts per second is at most (GPUs x peak FLOPs x MFU) / (2 x params x prompt tokens).
- Utilisation target from the p95 TTFT SLO; 60-75 percent of the binding resource is common for interactive traffic.
- Burstiness: size for peak minutes, not daily averages, then decide whether to queue, shed or scale for the bursts.
- Cost per 1M tokens = GPU dollars per hour / (tokens per hour at the planned utilisation); quote input and output separately or with a stated ratio.
Autoscaling signals compared
| Signal | Reacts to | Problem |
|---|---|---|
| Average GPU utilisation | Busy kernels | Near 100 percent at any load with continuous batching; lags tail latency |
| Requests per second | Arrivals | Ignores prompt and output length mix |
| Queue depth or waiting requests | Backlog | Good leading signal for TTFT |
| KV-cache usage | Memory pressure | Good leading signal for preemptions |
| p95 TTFT | The SLO itself | Lags; use as guardrail, not sole trigger |
GPU cold starts (pulling tens of GBs of weights) take minutes, so scale on leading signals and keep warm headroom.
Decision chart: what limits capacity?
Decisions
- 1
D1
- nextMemory-bound: FP8 KV, weight quantization, more GPUs per replica (TP)
- nextStep 2: is the prompt-to-output token ratio high (long prompts, short answers)?
- 2
Memory-bound: FP8 KV, weight quantization, more GPUs per replica (TP)
- nextRe-run the capacity model and the load test after every change
- ?
Step 2: is the prompt-to-output token ratio high (long prompts, short answers)?
- nextStep 3: do prompts share prefixes?
- nextStep 4: TPOT SLO met at the batch size needed for cost?
- ?
Step 3: do prompts share prefixes?
- nextPrefill-bound: prefix caching and cache-aware routing first
- nextPrefill-bound: chunked prefill, P/D disaggregation, FP8 compute
- 5
Prefill-bound: prefix caching and cache-aware routing first
- nextRe-run the capacity model and the load test after every change
- 6
Prefill-bound: chunked prefill, P/D disaggregation, FP8 compute
- nextRe-run the capacity model and the load test after every change
- ?
Step 4: TPOT SLO met at the batch size needed for cost?
- nextDecode-bound but healthy: raise batch until TPOT p95 nears the SLO
- nextDecode-bound: weight or KV quantization, speculative decoding, more replicas
- 8
Decode-bound but healthy: raise batch until TPOT p95 nears the SLO
- nextRe-run the capacity model and the load test after every change
- 9
Decode-bound: weight or KV quantization, speculative decoding, more replicas
- nextRe-run the capacity model and the load test after every change
- 10
Re-run the capacity model and the load test after every change
What happens if you choose otherwise
- Plan from average tokens/s: the fleet meets the average and fails p95 TTFT at every peak.
- Report TPOT only: a one-second stall every few hundred tokens disappears into the average while users see frozen streams.
- Autoscale on GPU utilisation: continuous batching keeps it high at any load, so scaling triggers late or never.
- Benchmark with uniform 512-in/128-out prompts: real prompts with long RAG context shift the bottleneck to prefill and halve capacity.
Pitfalls
- Ignoring the input/output token ratio when quoting $/1M tokens.
- Treating MFU and bandwidth efficiency as constants across batch sizes and sequence lengths; measure them on your engine.
- Forgetting N+1 capacity for replica failures and rolling deploys.
- Mixing interactive and batch traffic in one pool without priorities; batch jobs eat the headroom interactive SLOs need.
Interview Q&A
Define goodput and why it beats throughput.
Answer
Requests per second that meet all latency SLOs. Throughput can rise while SLO compliance collapses (huge batches); goodput measures only useful work.
TPOT vs ITL?
Answer
TPOT is an average per request; ITL is the distribution of gaps between tokens. Stalls from prefill interference or preemption show in ITL p99 and max but average away in TPOT.
Walk me through sizing a 70B deployment.
Answer
Fix SLOs and peak workload; compute memory for FP8 weights and KV per token to bound concurrency; compute prefill time per prompt and decode step time per batch; solve the steady state with Little's law for the max req/s per replica within SLOs; divide peak by that times headroom; add N+1; validate with a production-shaped load test; derive $/1M from GPU-hours and token volume.
Why not run GPUs at 95 percent utilisation?
Answer
Queueing delay grows sharply as utilisation approaches 100 percent; with random arrivals, p95 TTFT explodes, so the SLO sets a lower target.
What would you autoscale on?
Answer
Queue depth and KV-cache usage as leading signals, p95 TTFT as a guardrail, with warm headroom because model cold starts take minutes.
How was KV per token computed for the 70B example?
Answer
2 x 80 layers x 8 KV heads x 128 dims x 1 byte = 160 KiB, so memory left after weights holds about 600 concurrent sequences.
Why report input and output tokens separately?
Answer
Prefill and decode cost different resources.
Why keep warm headroom when autoscaling?
Answer
GPU cold starts pull tens of GBs of weights and take minutes, so scale on leading signals.
Check yourself
For one service, write its p95 TTFT and TPOT SLOs, peak req/s, prompt and output lengths. Estimate KV per token, the prefill share at peak, and the utilisation target your TTFT SLO allows.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Prefill vs Decode & KV Cache Mechanics, Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes, Speculative Decoding, Inference Parallelism — Tensor Parallel, Expert Parallel & Prefill/Decode Disaggregation, Histogram vs Average Latency SLIs, SLOs, Error Budgets & Distributed Tracing, Load Shedding & Admission Control — Drop Early, Protect the Core, LLM Token Streaming vs Parallel Decisions — SSE, TTFT & When Not to Stream, HPA, VPA & Autoscaling Gotchas — Metrics, Stabilization & Thrash, Rate Limiter for an LLM Gateway - LLD Spec & Concepts.
Go Deeper
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving
- NVIDIA NIM LLM benchmarking metrics
- Efficient Memory Management for LLM Serving with PagedAttention (vLLM)
- Sarathi-Serve: Taming Throughput-Latency Tradeoff with chunked prefills
- Splitwise: Efficient Generative LLM Inference Using Phase Splitting
- Google SRE Book: Service Level Objectives