LLM Inference Runtime
Studies in this cluster, in series order. Each one keeps its own URL.
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
LLM Inference Runtime
6 studies- 1.LLM Inference Runtime — Prefill, KV, Batching & ParallelismSenior interviews fail when they name a serving product before the bottleneck. This hub is the mechanism ladder under the live LLM Inference Engines cluster: prefill versus decode, KV bytes, PagedAttention, continuous batching, speculative decode, then serving tensor parallel, expert parallel, and prefill/decode disaggregation.
- 2.Prefill vs Decode & KV Cache MechanicsTime-to-first-token and tokens per second split because prefill and decode are different kernels. Prefill burns FLOPs on the prompt. Decode streams the KV cache from memory for one new token. This lesson is the byte math and the autoregressive loop those metrics sit on.
- 3.PagedAttention & Continuous BatchingPagedAttention and continuous batching are the default pair behind modern open serving. Block tables make KV virtual memory. An iteration-level scheduler admits and finishes requests without waiting for a static batch. RadixAttention is a different sharing layer and stays in the SGLang study.
- 4.Speculative DecodingWhen decode is memory-bandwidth bound, speculative decoding spends spare FLOPs to cut sequential target steps. A draft proposes several tokens and the target verifies them in one forward. Acceptance rate decides whether you win or waste the step.
- 5.Inference Parallelism — Tensor Parallel, Expert Parallel & Prefill/Decode DisaggregationServing scale-out is not the training mesh. This lesson is inference only: tensor parallel shards weights and KV and pays a collective on the decode path, expert parallel routes MoE tokens with an all-to-all, and prefill/decode disaggregation splits pools when TTFT and tokens per second fight.
- 6.vLLM vs SGLang — Runtime Choice (comparative)vLLM and SGLang share continuous batching and differ in the headline cache: paging versus a radix prefix tree. This page is a thin choice matrix. Prefix-tree internals, TensorRT engine builds, and playbook-length metric arithmetic stay in their own studies.