LLM Inference & Distributed Training
Studies in this cluster, in series order. Each one keeps its own URL.
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
LLM Inference & Distributed Training
6 studies- 1.LLM Inference Engines & Distributed Training — SGLang, TensorRT-LLM, ONNX & PyTorch ParallelismSenior SWE/ML interviews ask you to pick a serving stack and a parallelism strategy under latency SLOs, cost per token, hardware mix, and team maturity. Mis-picking SGLang vs TensorRT-LLM vs ONNX Runtime vs raw PyTorch — or confusing DDP with FSDP — produces symptoms that look like “the model is slow” but are architectural. This hub maps the cluster.
- 2.SGLang — RadixAttention, Continuous Batching & Structured GenerationMost LLM serving pain is the runtime, not the model. Agents, chat, RAG, and multi-turn loops resend long shared prefixes. SGLang’s bet is prefix reuse via RadixAttention, plus continuous batching and a constrained-decoding API. This lesson is the runtime layer — it does not re-teach JSON Schema or CFGs.
- 3.TensorRT-LLM — Engine Build, In-Flight Batching & QuantizationTensorRT-LLM compiles a model into a serialized TensorRT engine bound to one GPU SKU, one parallelism layout, and one quantization recipe. That buys NVIDIA latency and lock-in. This lesson covers convert → build → run, in-flight batching, FP8/INT4/INT8, multi-GPU topology, and the rebuild tax.
- 4.ONNX & ONNX Runtime — Export, Graph Optimizers & Execution ProvidersONNX is the interchange format; ONNX Runtime executes and rewrites the graph. Export cleanly, pick the right opset, and order Execution Providers — or you get numerical drift, unsupported-op failures, or a GPU that sits idle while nodes fall back to CPU. This lesson is the portable path, not an LLM-server lesson.
- 5.Distributed PyTorch — DDP, FSDP, Tensor & Pipeline ParallelismPast ~7B parameters, multi-GPU training failure modes change: rank-only OOM, unscaled LR, NCCL hangs from mismatched collectives, TP across a slow fabric. This lesson is the training mental model — DDP, ZeRO/FSDP, TP, PP, 3D grids — and why you should not import those intuitions into serving.
- 6.Comparative Playbook — Serving Stack Choice, Metrics & Failure ModesThe serving stack is where model quality meets a bill and a p99. Interviews test TTFT vs TPS, KV capacity math, quantization cliffs, and a defensible matrix — not “do you know vLLM.” This playbook ties the cluster together.