AI / ML
Topic, then cluster, then study. Recently added is the short list at the top.
Recently added
Show more- 6.Distillation & LLM Evals - Teacher-Student Training, Benchmarks, LLM-as-Judge, Regression Gates & ContaminationSequence-level vs logit-level vs on-policy distillation and reasoning-trace distillation; benchmarks, LLM-as-judge biases, regression gates with per-slice thresholds and CIs, contamination checks; runnable distill/contamination + regression gate.
- 1.LLM Inference in Production - Quantization, Prompt Caching, Multi-LoRA Serving & Capacity PlanningHub: fleet-level serving levers beyond the existing runtime pages; which lever (quantization, prefix caching, multi-LoRA, batching/capacity) moves TTFT, TPOT, memory or $/1M; runnable lever simulation + symptom-to-lever picker; triage chart.
- 5.LLM Latency, Cost & Capacity Planning - TTFT, TPOT/ITL, Goodput, $/1M Tokens & GPU SizingTTFT vs TPOT vs ITL, goodput, percentiles from histograms; worked 70B FP8 fleet sizing (KV memory, prefill share, Little's law, headroom, $/1M); Erlang C TTFT hockey stick; autoscaling signals compared; runnable capacity plan + TTFT calculator.
- 1.LLM Post-Training - From Base Model to Assistant: SFT, Preference Tuning, PEFT, Distillation & EvalsHub: base model -> assistant pipeline (SFT, RLHF/PPO, DPO-family and GRPO, LoRA/QLoRA, distillation, eval gate); techniques compared by signal needed, cost and failure mode; runnable pipeline toy + technique picker; decision chart.
- 2.LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offsRTN and per-tensor/channel/group scales, outliers, WxAy notation; GPTQ vs AWQ (W4A16), SmoothQuant (W8A8), FP8 E4M3; calibration, accuracy checks, when it helps TTFT vs only TPOT; runnable quant basics, GPTQ/AWQ toy, format math.
- 5.LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter MergingLoRA math (W + alpha/r BA), rank/alpha/target modules, rsLoRA/DoRA notes, QLoRA NF4 + double quantization + paged optimizers, merging vs keeping adapters, when PEFT is the wrong choice; runnable LoRA/NF4 demo + memory budget calculator.
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
Causal ML & Double Machine Learning
6 studies- 1.Causal ML & Double Machine Learning — From Association to EffectHub: correlation is not causation. Defend ATE, ATT, and CATE first, then identification, then estimators. Predictive ML is the wrong tool for treatment decisions; this cluster maps graphs, DML, CATE forests, causal time series, and health ethics.
- 2.Causal Graphs & Identification — DAGs, Confounders & BackdoorDAGs make identification inspectable before any DML fit. Distinguish confounders, colliders, and mediators; apply backdoor and positivity; name unmeasured confounding and confounding by indication.
- 3.Double Machine Learning — Nuisance Models, Orthogonalization & Cross-FittingChernozhukov DML: residual-on-residual / orthogonal scores, flexible ML nuisances, and K-fold cross-fitting so a low-dimensional ATE stays valid. DML does not invent identification — it beats naive ML-on-treatment when the backdoor set is already right.
- 4.Heterogeneous Treatment Effects — Causal Forests & Personalized MedicineCATE vs ATE: who benefits, not only whether the average helps. Causal forests and metalearners with honest splitting; personalized treatment only under overlap, multiplicity control, and validation — without overclaiming precision medicine.
- 5.Causal Time Series — ITS, Synthetic Control & Diff-in-Diff Over TimeInterrupted time series, synthetic control, and DiD over time estimate counterfactuals after a shock. Parallel trends and pre-fit beat a low forecast MAPE. Prediction under status quo is a different question — do not re-teach ARIMA, Prophet, or GBM.
- 6.Causal ML in Health — Outcomes, Bias, Ethics & ValidationHealth causal ML is a safety rail: endpoints, selection bias, confounding by indication, fairness, and validation before acting on an estimate. The bravest senior answer is often not deploying a CATE yet.
Choosing ML Algorithms
6 studies- 1.Choosing ML Algorithms — Problem Shape to Model FamilySenior interviewers do not want an algorithm encyclopedia. They want you to map problem shape to model family under constraints: tabular vs sequence, interpretability vs accuracy, latency, data size, and drift. Start with a baseline; ship the simplest model that meets the metric and the ops budget.
- 2.Decision Trees — Splits, Interpretability & When They FailA tree greedily partitions feature space into axis-aligned regions. That makes it readable — and brittle. Defend when a single tree is enough (debug, rules, small data, compliance) versus when it underfits or overfits, and why ensembles exist.
- 3.Random Forests & Bagging — Variance Reduction & Feature ImportanceA senior answer for “why random forests?” is variance reduction via averaging decorrelated trees — not “it is an ensemble.” Explain bootstrap aggregating, feature bagging, OOB error, and why impurity-based importance lies when features are correlated.
- 4.Gradient Boosted Trees — XGBoost / LightGBM / CatBoost Tradeoffs“Just use XGBoost” is not a senior answer. Explain sequential residual fitting, the learning-rate × n_estimators tradeoff, early stopping, and honest differences among XGBoost, LightGBM, and CatBoost — then defend when GBM beats RF, and when calibration or latency make RF or linear wiser.
- 5.Time Series Forecasting — Classical, ML Features & Deep SequencesTell a truly temporal problem from a table that happens to have a timestamp. Wrong choice means leakage, optimistic MAPE, and models that die after a holiday. Map classical ARIMA/ETS/Prophet, trees on lags, and deep sequences — with walk-forward validation. Prediction is not a counterfactual.
- 6.Algorithm Selection Playbook — Metrics, Baselines & Failure ModesChoosing XGBoost is easy; defending the full selection loop is not. Start with baselines, pick metrics that match business cost, handle imbalance, run a leakage checklist, know when not to use deep learning, and design for retrain cadence, drift, and compliance explainability.
LLM Inference & Distributed Training
6 studies- 1.LLM Inference Engines & Distributed Training — SGLang, TensorRT-LLM, ONNX & PyTorch ParallelismSenior SWE/ML interviews ask you to pick a serving stack and a parallelism strategy under latency SLOs, cost per token, hardware mix, and team maturity. Mis-picking SGLang vs TensorRT-LLM vs ONNX Runtime vs raw PyTorch — or confusing DDP with FSDP — produces symptoms that look like “the model is slow” but are architectural. This hub maps the cluster.
- 2.SGLang — RadixAttention, Continuous Batching & Structured GenerationMost LLM serving pain is the runtime, not the model. Agents, chat, RAG, and multi-turn loops resend long shared prefixes. SGLang’s bet is prefix reuse via RadixAttention, plus continuous batching and a constrained-decoding API. This lesson is the runtime layer — it does not re-teach JSON Schema or CFGs.
- 3.TensorRT-LLM — Engine Build, In-Flight Batching & QuantizationTensorRT-LLM compiles a model into a serialized TensorRT engine bound to one GPU SKU, one parallelism layout, and one quantization recipe. That buys NVIDIA latency and lock-in. This lesson covers convert → build → run, in-flight batching, FP8/INT4/INT8, multi-GPU topology, and the rebuild tax.
- 4.ONNX & ONNX Runtime — Export, Graph Optimizers & Execution ProvidersONNX is the interchange format; ONNX Runtime executes and rewrites the graph. Export cleanly, pick the right opset, and order Execution Providers — or you get numerical drift, unsupported-op failures, or a GPU that sits idle while nodes fall back to CPU. This lesson is the portable path, not an LLM-server lesson.
- 5.Distributed PyTorch — DDP, FSDP, Tensor & Pipeline ParallelismPast ~7B parameters, multi-GPU training failure modes change: rank-only OOM, unscaled LR, NCCL hangs from mismatched collectives, TP across a slow fabric. This lesson is the training mental model — DDP, ZeRO/FSDP, TP, PP, 3D grids — and why you should not import those intuitions into serving.
- 6.Comparative Playbook — Serving Stack Choice, Metrics & Failure ModesThe serving stack is where model quality meets a bill and a p99. Interviews test TTFT vs TPS, KV capacity math, quantization cliffs, and a defensible matrix — not “do you know vLLM.” This playbook ties the cluster together.
LLM Inference in Production
5 studies- 1.LLM Inference in Production - Quantization, Prompt Caching, Multi-LoRA Serving & Capacity PlanningHub: fleet-level serving levers beyond the existing runtime pages; which lever (quantization, prefix caching, multi-LoRA, batching/capacity) moves TTFT, TPOT, memory or $/1M; runnable lever simulation + symptom-to-lever picker; triage chart.
- 2.LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offsRTN and per-tensor/channel/group scales, outliers, WxAy notation; GPTQ vs AWQ (W4A16), SmoothQuant (W8A8), FP8 E4M3; calibration, accuracy checks, when it helps TTFT vs only TPOT; runnable quant basics, GPTQ/AWQ toy, format math.
- 3.Prefix & Prompt Caching Across Requests - Cache Keys, Routing, TTL & Hit-Rate EconomicsKV block hashing (vLLM APC) vs radix tree (SGLang); what goes into cache keys (adapter, tokenizer, images, salts); prefix-aware routing across replicas; provider prompt caching write premiums, read discounts, TTL; runnable fleet hit-rate sim + economics.
- 4.Multi-LoRA Serving - S-LoRA/Punica Batching, Adapter Hot-Swap & MemoryMerged-per-tenant vs merge-in-place vs multi-LoRA batching; Punica SGMV/BGMV gathered low-rank kernels; S-LoRA unified paging; GPU/host/disk adapter tiers, dynamic loading, adapter-affinity routing, rank caps; runnable batching + adapter cache sims.
- 5.LLM Latency, Cost & Capacity Planning - TTFT, TPOT/ITL, Goodput, $/1M Tokens & GPU SizingTTFT vs TPOT vs ITL, goodput, percentiles from histograms; worked 70B FP8 fleet sizing (KV memory, prefill share, Little's law, headroom, $/1M); Erlang C TTFT hockey stick; autoscaling signals compared; runnable capacity plan + TTFT calculator.
LLM Inference Runtime
6 studies- 1.LLM Inference Runtime — Prefill, KV, Batching & ParallelismSenior interviews fail when they name a serving product before the bottleneck. This hub is the mechanism ladder under the live LLM Inference Engines cluster: prefill versus decode, KV bytes, PagedAttention, continuous batching, speculative decode, then serving tensor parallel, expert parallel, and prefill/decode disaggregation.
- 2.Prefill vs Decode & KV Cache MechanicsTime-to-first-token and tokens per second split because prefill and decode are different kernels. Prefill burns FLOPs on the prompt. Decode streams the KV cache from memory for one new token. This lesson is the byte math and the autoregressive loop those metrics sit on.
- 3.PagedAttention & Continuous BatchingPagedAttention and continuous batching are the default pair behind modern open serving. Block tables make KV virtual memory. An iteration-level scheduler admits and finishes requests without waiting for a static batch. RadixAttention is a different sharing layer and stays in the SGLang study.
- 4.Speculative DecodingWhen decode is memory-bandwidth bound, speculative decoding spends spare FLOPs to cut sequential target steps. A draft proposes several tokens and the target verifies them in one forward. Acceptance rate decides whether you win or waste the step.
- 5.Inference Parallelism — Tensor Parallel, Expert Parallel & Prefill/Decode DisaggregationServing scale-out is not the training mesh. This lesson is inference only: tensor parallel shards weights and KV and pays a collective on the decode path, expert parallel routes MoE tokens with an all-to-all, and prefill/decode disaggregation splits pools when TTFT and tokens per second fight.
- 6.vLLM vs SGLang — Runtime Choice (comparative)vLLM and SGLang share continuous batching and differ in the headline cache: paging versus a radix prefix tree. This page is a thin choice matrix. Prefix-tree internals, TensorRT engine builds, and playbook-length metric arithmetic stay in their own studies.
LLM Post-Training
6 studies- 1.LLM Post-Training - From Base Model to Assistant: SFT, Preference Tuning, PEFT, Distillation & EvalsHub: base model -> assistant pipeline (SFT, RLHF/PPO, DPO-family and GRPO, LoRA/QLoRA, distillation, eval gate); techniques compared by signal needed, cost and failure mode; runnable pipeline toy + technique picker; decision chart.
- 2.Supervised Fine-Tuning (SFT) - Instruction Data, Chat Templates, Packing & Loss MaskingSFT data formats, chat templates and special tokens, assistant-only loss masking, packing with attention boundaries, truncation; data quality over volume; runnable masking + packing demos; template/masking bug failure paths.
- 3.Reward Models & RLHF with PPO - Bradley-Terry, KL Penalty, Critics & Reward HackingBradley-Terry reward models from pairwise preferences, PPO with KL penalty to the SFT reference, critic/GAE, four-models-in-memory cost, reward hacking and over-optimisation; runnable RM hacking demo + PPO step.
- 4.Preference Optimization - DPO, IPO, ORPO, SimPO, KTO & GRPO Compared with PPODPO derivation from the RLHF objective; IPO, KTO, ORPO, SimPO variants (reference-free, unpaired data); GRPO group-relative advantages with verifiable rewards; all compared with PPO by signal, cost, failure mode; runnable losses + GRPO advantages.
- 5.LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter MergingLoRA math (W + alpha/r BA), rank/alpha/target modules, rsLoRA/DoRA notes, QLoRA NF4 + double quantization + paged optimizers, merging vs keeping adapters, when PEFT is the wrong choice; runnable LoRA/NF4 demo + memory budget calculator.
- 6.Distillation & LLM Evals - Teacher-Student Training, Benchmarks, LLM-as-Judge, Regression Gates & ContaminationSequence-level vs logit-level vs on-policy distillation and reasoning-trace distillation; benchmarks, LLM-as-judge biases, regression gates with per-slice thresholds and CIs, contamination checks; runnable distill/contamination + regression gate.
Production AI Agents
6 studies- 1.Production AI Agents — Architecture, Scaling, Caching & ReliabilityA production AI agent is a policy loop: model, tools, memory, and guardrails, with durable state, cost and latency SLOs, and cache layers. It is not a chat box glued to one completion. This hub maps architecture, memory, scale-out, caching, and reliability, and only cross-links the RAG, structured-outputs, and inference hubs.
- 2.Agent Architecture — Planner/Executor, Tool Registry & Control LoopsProduction agents need an explicit control architecture: who plans, who executes, how tools are registered and validated, and where humans gate irreversible actions. Compare ReAct, plan-then-act, and a graph, and point at structured outputs for argument shape without re-teaching constrained decoding.
- 3.Agent Memory — Working Context, Long-term Store, Compaction & SessionsAgent memory is a budgeted systems problem. Working context fits in the window. Episodes and semantic facts live outside it. Compaction must keep goals and commitments and must not leak PII across sessions. Retrieval of documents is a different hub.
- 4.Scaling Agents — Concurrency, Queues, Fan-out, Rate Limits & Cost BudgetsScaling agents is a distributed-systems problem: horizontal workers, job queues, per-tenant quotas, backpressure, and idempotent tool side effects. Every parallel tool call is a budgeted fork-join. Fan-out multiplies capability and cost together.
- 5.Caching for Agents — Prompt Cache, KV Reuse, Semantic Cache & Tool Result CacheThe wrong agent cache is a correctness bug. Separate provider prompt cache, runtime KV reuse, semantic response cache, and tool-result memoization. This page is the application policy. PagedAttention and the KV layout stay on the inference hub.
- 6.Agent Reliability — Evals, Guardrails, Tracing & Failure ModesAgents fail differently from chatbots: stuck loops, poison tool output, partial side effects, and trajectory drift. Reliability means grading whole runs, sandboxing tools, circuit breakers, and span-level traces. A fluent final string is not a pass.
RAG & Vector Databases
6 studies- 1.RAG & Vector Databases — Retrieval, Embeddings & GroundingRAG grounds an LLM on your documents at query time: embed, retrieve, assemble context, and generate with citations. This hub is the decision map for when retrieval beats fine-tuning or long context, and how to judge grounding.
- 2.Embeddings & Similarity — Dense Vectors, Metrics & Chunking BasicsEmbeddings map text to dense vectors so similar meaning lands nearby. The usable part is the metric, the model, and the chunk size. A strong embedder with a 4k-token mush window still retrieves mush.
- 3.Vector Indexes — HNSW, IVF & Product Quantization TradeoffsBrute-force top-k dies at millions of vectors. HNSW, IVF, and product quantization trade a measured amount of recall for latency and memory. Pick them the way you pick a B-tree versus a hash index: with a recall curve and a p99.
- 4.Hybrid Retrieval — BM25 + Vectors, Reranking & Metadata FiltersProduction RAG rarely ships vectors alone. Hybrid retrieval fuses BM25 with dense similarity, reranks a shortlist, and applies metadata filters for tenant, ACL, and time before the model ever sees a chunk.
- 5.RAG Context Assembly — Windows, Parent-Child Chunks & CitationsRetrieval returns candidates. Context assembly decides what the model actually sees: a token budget, deduped parents, neighbor windows, and citation ids. A perfect top-k still fails if you dump it raw into the prompt.
- 6.RAG Failure Modes — Hallucination, Stale Indexes, Evals & GroundingRAG fails loudly, with an empty answer, or quietly, with a confident citation that does not support the claim. This lesson is the catalog: hallucination, stale indexes, tenant leaks, and the evals that catch them before a doc edit rots the demo.
Structured outputs
6 studies- 1.Structured Outputs / Constrained DecodingStrict JSON Schema vs JSON mode; CFG token masking; required fields + additionalProperties:false; Pydantic/Zod; refusals/truncation still break validity.
- 2.JSON Schema Strictness for Structured OutputsStrict structured outputs need a closed schema: root object, every property required, additionalProperties false, provider-supported subset, schema_version, null for absence.
- 3.Tool/Function Calling vs Structured OutputsTools execute side effects and fetch live data; structured outputs constrain a final JSON contract; agent loops mix both. Pick by side effects, latency, and security — not by habit.
- 4.Grammars & CFG Constrained DecodingCompile a CFG/GBNF to an automaton, mask illegal next tokens, keep a parse stack. Same idea as JSON Schema SO, but you can constrain SQL, arithmetic, or custom DSLs.
- 5.Validation & Repair Loops for LLM OutputsPrefer constrained decoding for syntax. Use a bounded validate→feedback→regenerate loop for business rules or legacy models. Never repair refusals; never execute tools on invalid JSON.
- 6.Streaming Structured OutputStreaming tokens of JSON improves UX but partial JSON is invalid. Incremental parsers paint completed fields; commit side effects only after final schema validation. Constrained decoding still applies per token.
System One & Jev
6 studies- 1.System One Models & Jev — Fast Structured Decisions for SoftwareChat models optimize for preferred strings. Software needs typed, probabilistic decisions it can branch on without parsing prose. System One Models (TypeSafe, announced 2026-09-15) take unstructured state in and emit typed probabilistic decisions. Jev is the first public System One model — this hub maps the cluster.
- 2.Choice, Score & Noul — State, Parallel Questions & Typed AnswersSystem One usefulness lives in three primitives. Mis-picking the type is a common interview fail: using a Noul when you need ordered levels, or a free-form LLM parse when a closed Choice would do. This lesson teaches state, question IDs, criteria, parallel evaluation, and answer shapes — with sandbox mocks and no API keys.
- 3.Confidence-Gated Routing, Composite Scoring & Workflow DecompositionTyped answers alone do not make automation safe. Confidence (Choice/Score) and noul magnitude (yes/no) are the second axis: what vs whether to act. This lesson teaches stake-scaled thresholds, composite scores owned in code, and speculative fan-out — without re-teaching Structured Outputs validation loops.
- 4.Generators, yield & yield* — Composing Streams and Decision WorkflowsTwo composition styles collide in AI backends: sequential token streams via generators/yield, and one-shot parallel decision round-trips (System One). Interviews expect you to implement async generators for LLM chunks and to compose typed decision steps without confusing them with token streaming.
- 5.LLM Token Streaming vs Parallel Decisions — SSE, TTFT & When Not to StreamStaff interviews split “stream tokens for UX” from “await a structured decision.” Mixing them causes bad architectures: streaming a classifier, or blocking the UI for a paragraph that should have streamed. This lesson covers TTFT, SSE/chunked tokens, cancellation, backpressure, and when a System One parallel call is enough.
- 6.Hybrid Architecture — System One for Route/Guardrail, LLMs for ProseProduction systems rarely pick one model class. The winning shape: System One (or equivalent classifiers) for route, score, and guardrail; generative LLMs for explanations, drafts, and open-ended tools; deterministic code as the source of truth. This capstone wires primitives, confidence, generators, and streaming into an interview-ready architecture.