LLM Inference in Production
Studies in this cluster, in series order. Each one keeps its own URL.
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
LLM Inference in Production
5 studies- 1.LLM Inference in Production - Quantization, Prompt Caching, Multi-LoRA Serving & Capacity PlanningHub: fleet-level serving levers beyond the existing runtime pages; which lever (quantization, prefix caching, multi-LoRA, batching/capacity) moves TTFT, TPOT, memory or $/1M; runnable lever simulation + symptom-to-lever picker; triage chart.
- 2.LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offsRTN and per-tensor/channel/group scales, outliers, WxAy notation; GPTQ vs AWQ (W4A16), SmoothQuant (W8A8), FP8 E4M3; calibration, accuracy checks, when it helps TTFT vs only TPOT; runnable quant basics, GPTQ/AWQ toy, format math.
- 3.Prefix & Prompt Caching Across Requests - Cache Keys, Routing, TTL & Hit-Rate EconomicsKV block hashing (vLLM APC) vs radix tree (SGLang); what goes into cache keys (adapter, tokenizer, images, salts); prefix-aware routing across replicas; provider prompt caching write premiums, read discounts, TTL; runnable fleet hit-rate sim + economics.
- 4.Multi-LoRA Serving - S-LoRA/Punica Batching, Adapter Hot-Swap & MemoryMerged-per-tenant vs merge-in-place vs multi-LoRA batching; Punica SGMV/BGMV gathered low-rank kernels; S-LoRA unified paging; GPU/host/disk adapter tiers, dynamic loading, adapter-affinity routing, rank caps; runnable batching + adapter cache sims.
- 5.LLM Latency, Cost & Capacity Planning - TTFT, TPOT/ITL, Goodput, $/1M Tokens & GPU SizingTTFT vs TPOT vs ITL, goodput, percentiles from histograms; worked 70B FP8 fleet sizing (KV memory, prefill share, Little's law, headroom, $/1M); Erlang C TTFT hockey stick; autoscaling signals compared; runnable capacity plan + TTFT calculator.