TopicsAI / ML
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
Common tags: rag, embeddings, evals, serving
- AI / ML
Supervised Fine-Tuning (SFT) - Instruction Data, Chat Templates, Packing & Loss Masking
Cluster · LLM Post-Training
SFT data formats, chat templates and special tokens, assistant-only loss masking, packing with attention boundaries, truncation; data quality over volume; runnable masking + packing demos; template/masking bug failure paths.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
Reward Models & RLHF with PPO - Bradley-Terry, KL Penalty, Critics & Reward Hacking
Cluster · LLM Post-Training
Bradley-Terry reward models from pairwise preferences, PPO with KL penalty to the SFT reference, critic/GAE, four-models-in-memory cost, reward hacking and over-optimisation; runnable RM hacking demo + PPO step.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
Prefix & Prompt Caching Across Requests - Cache Keys, Routing, TTL & Hit-Rate Economics
Cluster · LLM Inference in Production
KV block hashing (vLLM APC) vs radix tree (SGLang); what goes into cache keys (adapter, tokenizer, images, salts); prefix-aware routing across replicas; provider prompt caching write premiums, read discounts, TTL; runnable fleet hit-rate sim + economics.
Open study →- ai-ml
- llm
- inference
- serving
- quantization
- gptq
- awq
- fp8
- prefix-caching
- prompt-caching
- multi-lora
- s-lora
- capacity-planning
- ttft
- goodput
- gpu-sizing
- interview
- AI / ML
Preference Optimization - DPO, IPO, ORPO, SimPO, KTO & GRPO Compared with PPO
Cluster · LLM Post-Training
DPO derivation from the RLHF objective; IPO, KTO, ORPO, SimPO variants (reference-free, unpaired data); GRPO group-relative advantages with verifiable rewards; all compared with PPO by signal, cost, failure mode; runnable losses + GRPO advantages.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
Multi-LoRA Serving - S-LoRA/Punica Batching, Adapter Hot-Swap & Memory
Cluster · LLM Inference in Production
Merged-per-tenant vs merge-in-place vs multi-LoRA batching; Punica SGMV/BGMV gathered low-rank kernels; S-LoRA unified paging; GPU/host/disk adapter tiers, dynamic loading, adapter-affinity routing, rank caps; runnable batching + adapter cache sims.
Open study →- ai-ml
- llm
- inference
- serving
- quantization
- gptq
- awq
- fp8
- prefix-caching
- prompt-caching
- multi-lora
- s-lora
- capacity-planning
- ttft
- goodput
- gpu-sizing
- interview
- AI / ML
LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter Merging
Cluster · LLM Post-Training
LoRA math (W + alpha/r BA), rank/alpha/target modules, rsLoRA/DoRA notes, QLoRA NF4 + double quantization + paged optimizers, merging vs keeping adapters, when PEFT is the wrong choice; runnable LoRA/NF4 demo + memory budget calculator.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offs
Cluster · LLM Inference in Production
RTN and per-tensor/channel/group scales, outliers, WxAy notation; GPTQ vs AWQ (W4A16), SmoothQuant (W8A8), FP8 E4M3; calibration, accuracy checks, when it helps TTFT vs only TPOT; runnable quant basics, GPTQ/AWQ toy, format math.
Open study →- ai-ml
- llm
- inference
- serving
- quantization
- gptq
- awq
- fp8
- prefix-caching
- prompt-caching
- multi-lora
- s-lora
- capacity-planning
- ttft
- goodput
- gpu-sizing
- interview
- AI / ML
LLM Post-Training - From Base Model to Assistant: SFT, Preference Tuning, PEFT, Distillation & Evals
Cluster · LLM Post-Training
Hub: base model -> assistant pipeline (SFT, RLHF/PPO, DPO-family and GRPO, LoRA/QLoRA, distillation, eval gate); techniques compared by signal needed, cost and failure mode; runnable pipeline toy + technique picker; decision chart.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
LLM Latency, Cost & Capacity Planning - TTFT, TPOT/ITL, Goodput, $/1M Tokens & GPU Sizing
Cluster · LLM Inference in Production
TTFT vs TPOT vs ITL, goodput, percentiles from histograms; worked 70B FP8 fleet sizing (KV memory, prefill share, Little's law, headroom, $/1M); Erlang C TTFT hockey stick; autoscaling signals compared; runnable capacity plan + TTFT calculator.
Open study →- ai-ml
- llm
- inference
- serving
- quantization
- gptq
- awq
- fp8
- prefix-caching
- prompt-caching
- multi-lora
- s-lora
- capacity-planning
- ttft
- goodput
- gpu-sizing
- interview
- AI / ML
LLM Inference in Production - Quantization, Prompt Caching, Multi-LoRA Serving & Capacity Planning
Cluster · LLM Inference in Production
Hub: fleet-level serving levers beyond the existing runtime pages; which lever (quantization, prefix caching, multi-LoRA, batching/capacity) moves TTFT, TPOT, memory or $/1M; runnable lever simulation + symptom-to-lever picker; triage chart.
Open study →- ai-ml
- llm
- inference
- serving
- quantization
- gptq
- awq
- fp8
- prefix-caching
- prompt-caching
- multi-lora
- s-lora
- capacity-planning
- ttft
- goodput
- gpu-sizing
- interview
- AI / ML
Distillation & LLM Evals - Teacher-Student Training, Benchmarks, LLM-as-Judge, Regression Gates & Contamination
Cluster · LLM Post-Training
Sequence-level vs logit-level vs on-policy distillation and reasoning-trace distillation; benchmarks, LLM-as-judge biases, regression gates with per-slice thresholds and CIs, contamination checks; runnable distill/contamination + regression gate.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
Production AI Agents — Architecture, Scaling, Caching & Reliability
Cluster · Production AI Agents
A production AI agent is a policy loop: model, tools, memory, and guardrails, with durable state, cost and latency SLOs, and cache layers. It is not a chat box glued to one completion. This hub maps architecture, memory, scale-out, caching, and reliability, and only cross-links the RAG, structured-outputs, and inference hubs.
Open study →- ai-ml
- agents
- orchestration
- scaling
- caching
- reliability
- interview
- AI / ML
Scaling Agents — Concurrency, Queues, Fan-out, Rate Limits & Cost Budgets
Cluster · Production AI Agents
Scaling agents is a distributed-systems problem: horizontal workers, job queues, per-tenant quotas, backpressure, and idempotent tool side effects. Every parallel tool call is a budgeted fork-join. Fan-out multiplies capability and cost together.
Open study →- ai-ml
- agents
- orchestration
- scaling
- caching
- reliability
- interview
- AI / ML
Agent Reliability — Evals, Guardrails, Tracing & Failure Modes
Cluster · Production AI Agents
Agents fail differently from chatbots: stuck loops, poison tool output, partial side effects, and trajectory drift. Reliability means grading whole runs, sandboxing tools, circuit breakers, and span-level traces. A fluent final string is not a pass.
Open study →- ai-ml
- agents
- orchestration
- scaling
- caching
- reliability
- interview
- AI / ML
Agent Memory — Working Context, Long-term Store, Compaction & Sessions
Cluster · Production AI Agents
Agent memory is a budgeted systems problem. Working context fits in the window. Episodes and semantic facts live outside it. Compaction must keep goals and commitments and must not leak PII across sessions. Retrieval of documents is a different hub.
Open study →- ai-ml
- agents
- orchestration
- scaling
- caching
- reliability
- interview
- AI / ML
Caching for Agents — Prompt Cache, KV Reuse, Semantic Cache & Tool Result Cache
Cluster · Production AI Agents
The wrong agent cache is a correctness bug. Separate provider prompt cache, runtime KV reuse, semantic response cache, and tool-result memoization. This page is the application policy. PagedAttention and the KV layout stay on the inference hub.
Open study →- ai-ml
- agents
- orchestration
- scaling
- caching
- reliability
- interview
- AI / ML
Agent Architecture — Planner/Executor, Tool Registry & Control Loops
Cluster · Production AI Agents
Production agents need an explicit control architecture: who plans, who executes, how tools are registered and validated, and where humans gate irreversible actions. Compare ReAct, plan-then-act, and a graph, and point at structured outputs for argument shape without re-teaching constrained decoding.
Open study →- ai-ml
- agents
- orchestration
- scaling
- caching
- reliability
- interview
- AI / ML
Vector Indexes — HNSW, IVF & Product Quantization Tradeoffs
Cluster · RAG & Vector Databases
Brute-force top-k dies at millions of vectors. HNSW, IVF, and product quantization trade a measured amount of recall for latency and memory. Pick them the way you pick a B-tree versus a hash index: with a recall curve and a p99.
Open study →- ai-ml
- rag
- embeddings
- vector-databases
- retrieval
- grounding
- interview
- AI / ML
RAG & Vector Databases — Retrieval, Embeddings & Grounding
Cluster · RAG & Vector Databases
RAG grounds an LLM on your documents at query time: embed, retrieve, assemble context, and generate with citations. This hub is the decision map for when retrieval beats fine-tuning or long context, and how to judge grounding.
Open study →- ai-ml
- rag
- embeddings
- vector-databases
- retrieval
- grounding
- interview
- AI / ML
RAG Failure Modes — Hallucination, Stale Indexes, Evals & Grounding
Cluster · RAG & Vector Databases
RAG fails loudly, with an empty answer, or quietly, with a confident citation that does not support the claim. This lesson is the catalog: hallucination, stale indexes, tenant leaks, and the evals that catch them before a doc edit rots the demo.
Open study →- ai-ml
- rag
- embeddings
- vector-databases
- retrieval
- grounding
- interview
- AI / ML
RAG Context Assembly — Windows, Parent-Child Chunks & Citations
Cluster · RAG & Vector Databases
Retrieval returns candidates. Context assembly decides what the model actually sees: a token budget, deduped parents, neighbor windows, and citation ids. A perfect top-k still fails if you dump it raw into the prompt.
Open study →- ai-ml
- rag
- embeddings
- vector-databases
- retrieval
- grounding
- interview
- AI / ML
Hybrid Retrieval — BM25 + Vectors, Reranking & Metadata Filters
Cluster · RAG & Vector Databases
Production RAG rarely ships vectors alone. Hybrid retrieval fuses BM25 with dense similarity, reranks a shortlist, and applies metadata filters for tenant, ACL, and time before the model ever sees a chunk.
Open study →- ai-ml
- rag
- embeddings
- vector-databases
- retrieval
- grounding
- interview
- AI / ML
Embeddings & Similarity — Dense Vectors, Metrics & Chunking Basics
Cluster · RAG & Vector Databases
Embeddings map text to dense vectors so similar meaning lands nearby. The usable part is the metric, the model, and the chunk size. A strong embedder with a 4k-token mush window still retrieves mush.
Open study →- ai-ml
- rag
- embeddings
- vector-databases
- retrieval
- grounding
- interview
- AI / ML
vLLM vs SGLang — Runtime Choice (comparative)
Cluster · LLM Inference Runtime
vLLM and SGLang share continuous batching and differ in the headline cache: paging versus a radix prefix tree. This page is a thin choice matrix. Prefix-tree internals, TensorRT engine builds, and playbook-length metric arithmetic stay in their own studies.
Open study →- ai-ml
- vllm
- sglang
- serving
- decision-matrix
- interview
- AI / ML
Speculative Decoding
Cluster · LLM Inference Runtime
When decode is memory-bandwidth bound, speculative decoding spends spare FLOPs to cut sequential target steps. A draft proposes several tokens and the target verifies them in one forward. Acceptance rate decides whether you win or waste the step.
Open study →- ai-ml
- speculative-decoding
- draft-model
- medusa
- eagle
- interview
- AI / ML
Prefill vs Decode & KV Cache Mechanics
Cluster · LLM Inference Runtime
Time-to-first-token and tokens per second split because prefill and decode are different kernels. Prefill burns FLOPs on the prompt. Decode streams the KV cache from memory for one new token. This lesson is the byte math and the autoregressive loop those metrics sit on.
Open study →- ai-ml
- prefill
- decode
- kv-cache
- ttft
- tpot
- interview
- AI / ML
PagedAttention & Continuous Batching
Cluster · LLM Inference Runtime
PagedAttention and continuous batching are the default pair behind modern open serving. Block tables make KV virtual memory. An iteration-level scheduler admits and finishes requests without waiting for a static batch. RadixAttention is a different sharing layer and stays in the SGLang study.
Open study →- ai-ml
- pagedattention
- continuous-batching
- vllm
- kv-cache
- interview
- AI / ML
LLM Inference Runtime — Prefill, KV, Batching & Parallelism
Cluster · LLM Inference Runtime
Senior interviews fail when they name a serving product before the bottleneck. This hub is the mechanism ladder under the live LLM Inference Engines cluster: prefill versus decode, KV bytes, PagedAttention, continuous batching, speculative decode, then serving tensor parallel, expert parallel, and prefill/decode disaggregation.
Open study →- ai-ml
- llm-inference
- kv-cache
- pagedattention
- continuous-batching
- speculative-decoding
- tensor-parallel
- expert-parallel
- vllm
- sglang
- interview
- AI / ML
Inference Parallelism — Tensor Parallel, Expert Parallel & Prefill/Decode Disaggregation
Cluster · LLM Inference Runtime
Serving scale-out is not the training mesh. This lesson is inference only: tensor parallel shards weights and KV and pays a collective on the decode path, expert parallel routes MoE tokens with an all-to-all, and prefill/decode disaggregation splits pools when TTFT and tokens per second fight.
Open study →- ai-ml
- tensor-parallel
- expert-parallel
- moe
- disaggregation
- interview
- AI / ML
TensorRT-LLM — Engine Build, In-Flight Batching & Quantization
Cluster · LLM Inference & Distributed Training
TensorRT-LLM compiles a model into a serialized TensorRT engine bound to one GPU SKU, one parallelism layout, and one quantization recipe. That buys NVIDIA latency and lock-in. This lesson covers convert → build → run, in-flight batching, FP8/INT4/INT8, multi-GPU topology, and the rebuild tax.
Open study →- ai-ml
- tensorrt-llm
- nvidia
- inference
- quantization
- fp8
- int4
- in-flight-batching
- paged-kv-cache
- tensor-parallel
- engine-build
- latency-optimization
- model-serving
- interview
- AI / ML
System One Models & Jev — Fast Structured Decisions for Software
Cluster · System One & Jev
Chat models optimize for preferred strings. Software needs typed, probabilistic decisions it can branch on without parsing prose. System One Models (TypeSafe, announced 2026-09-15) take unstructured state in and emit typed probabilistic decisions. Jev is the first public System One model — this hub maps the cluster.
Open study →- ai-ml
- system-one
- jev
- typesafe
- structured-decisions
- rlcd
- interview
- AI / ML
Confidence-Gated Routing, Composite Scoring & Workflow Decomposition
Cluster · System One & Jev
Typed answers alone do not make automation safe. Confidence (Choice/Score) and noul magnitude (yes/no) are the second axis: what vs whether to act. This lesson teaches stake-scaled thresholds, composite scores owned in code, and speculative fan-out — without re-teaching Structured Outputs validation loops.
Open study →- ai-ml
- confidence
- routing
- composite-scoring
- fan-out
- system-one
- interview
- AI / ML
SGLang — RadixAttention, Continuous Batching & Structured Generation
Cluster · LLM Inference & Distributed Training
Most LLM serving pain is the runtime, not the model. Agents, chat, RAG, and multi-turn loops resend long shared prefixes. SGLang’s bet is prefix reuse via RadixAttention, plus continuous batching and a constrained-decoding API. This lesson is the runtime layer — it does not re-teach JSON Schema or CFGs.
Open study →- ai-ml
- sglang
- radixattention
- prefix-cache
- continuous-batching
- structured-generation
- llm-serving
- openai-compatible
- inference-runtime
- kv-cache
- interview
- AI / ML
ONNX & ONNX Runtime — Export, Graph Optimizers & Execution Providers
Cluster · LLM Inference & Distributed Training
ONNX is the interchange format; ONNX Runtime executes and rewrites the graph. Export cleanly, pick the right opset, and order Execution Providers — or you get numerical drift, unsupported-op failures, or a GPU that sits idle while nodes fall back to CPU. This lesson is the portable path, not an LLM-server lesson.
Open study →- ai-ml
- onnx
- onnx-runtime
- model-export
- execution-providers
- graph-optimizers
- graph-optimization
- tensorrt
- coreml
- quantization
- deployment
- inference
- interview
- AI / ML
LLM Token Streaming vs Parallel Decisions — SSE, TTFT & When Not to Stream
Cluster · System One & Jev
Staff interviews split “stream tokens for UX” from “await a structured decision.” Mixing them causes bad architectures: streaming a classifier, or blocking the UI for a paragraph that should have streamed. This lesson covers TTFT, SSE/chunked tokens, cancellation, backpressure, and when a System One parallel call is enough.
Open study →- ai-ml
- llm-streaming
- sse
- ttft
- abortsignal
- system-one
- interview
- AI / ML
LLM Inference Engines & Distributed Training — SGLang, TensorRT-LLM, ONNX & PyTorch Parallelism
Cluster · LLM Inference & Distributed Training
Senior SWE/ML interviews ask you to pick a serving stack and a parallelism strategy under latency SLOs, cost per token, hardware mix, and team maturity. Mis-picking SGLang vs TensorRT-LLM vs ONNX Runtime vs raw PyTorch — or confusing DDP with FSDP — produces symptoms that look like “the model is slow” but are architectural. This hub maps the cluster.
Open study →- ai-ml
- llm
- llm-inference
- sglang
- tensorrt-llm
- onnx
- pytorch
- distributed-training
- ddp
- fsdp
- serving
- kv-cache
- continuous-batching
- quantization
- interview
- AI / ML
Choice, Score & Noul — State, Parallel Questions & Typed Answers
Cluster · System One & Jev
System One usefulness lives in three primitives. Mis-picking the type is a common interview fail: using a Noul when you need ordered levels, or a free-form LLM parse when a closed Choice would do. This lesson teaches state, question IDs, criteria, parallel evaluation, and answer shapes — with sandbox mocks and no API keys.
Open study →- ai-ml
- jev
- choice
- score
- noul
- primitives
- system-one
- interview
- AI / ML
Hybrid Architecture — System One for Route/Guardrail, LLMs for Prose
Cluster · System One & Jev
Production systems rarely pick one model class. The winning shape: System One (or equivalent classifiers) for route, score, and guardrail; generative LLMs for explanations, drafts, and open-ended tools; deterministic code as the source of truth. This capstone wires primitives, confidence, generators, and streaming into an interview-ready architecture.
Open study →- ai-ml
- hybrid
- guardrails
- routing
- llm
- system-one
- jev
- interview
- AI / ML
Generators, yield & yield* — Composing Streams and Decision Workflows
Cluster · System One & Jev
Two composition styles collide in AI backends: sequential token streams via generators/yield, and one-shot parallel decision round-trips (System One). Interviews expect you to implement async generators for LLM chunks and to compose typed decision steps without confusing them with token streaming.
Open study →- ai-ml
- generators
- yield
- async-generators
- effect
- streaming
- system-one
- interview
- AI / ML
Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism
Cluster · LLM Inference & Distributed Training
Past ~7B parameters, multi-GPU training failure modes change: rank-only OOM, unscaled LR, NCCL hangs from mismatched collectives, TP across a slow fabric. This lesson is the training mental model — DDP, ZeRO/FSDP, TP, PP, 3D grids — and why you should not import those intuitions into serving.
Open study →- ai-ml
- distributed-training
- ddp
- fsdp
- zero
- tensor-parallel
- pipeline-parallel
- nccl
- collectives
- pytorch
- 3d-parallelism
- device-mesh
- interview
- AI / ML
Comparative Playbook — Serving Stack Choice, Metrics & Failure Modes
Cluster · LLM Inference & Distributed Training
The serving stack is where model quality meets a bill and a p99. Interviews test TTFT vs TPS, KV capacity math, quantization cliffs, and a defensible matrix — not “do you know vLLM.” This playbook ties the cluster together.
Open study →- ai-ml
- serving
- inference
- ttft
- throughput
- kv-cache
- quantization
- sglang
- tensorrt-llm
- onnx-runtime
- pytorch
- capacity-planning
- cost-modeling
- llm-ops
- decision-matrix
- ops
- interview
- AI / ML
Time Series Forecasting — Classical, ML Features & Deep Sequences
Cluster · Choosing ML Algorithms
Tell a truly temporal problem from a table that happens to have a timestamp. Wrong choice means leakage, optimistic MAPE, and models that die after a holiday. Map classical ARIMA/ETS/Prophet, trees on lags, and deep sequences — with walk-forward validation. Prediction is not a counterfactual.
Open study →- ai-ml
- machine-learning
- ml
- interview
- time-series
- forecasting
- arima
- prophet
- walk-forward
- leakage
- AI / ML
Random Forests & Bagging — Variance Reduction & Feature Importance
Cluster · Choosing ML Algorithms
A senior answer for “why random forests?” is variance reduction via averaging decorrelated trees — not “it is an ensemble.” Explain bootstrap aggregating, feature bagging, OOB error, and why impurity-based importance lies when features are correlated.
Open study →- ai-ml
- machine-learning
- ml
- interview
- random-forest
- bagging
- oob
- feature-importance
- ensemble
- AI / ML
Algorithm Selection Playbook — Metrics, Baselines & Failure Modes
Cluster · Choosing ML Algorithms
Choosing XGBoost is easy; defending the full selection loop is not. Start with baselines, pick metrics that match business cost, handle imbalance, run a leakage checklist, know when not to use deep learning, and design for retrain cadence, drift, and compliance explainability.
Open study →- ai-ml
- machine-learning
- ml
- interview
- metrics
- baselines
- leakage
- calibration
- production-ml
- algorithm-selection
- AI / ML
Heterogeneous Treatment Effects — Causal Forests & Personalized Medicine
Cluster · Causal ML & Double Machine Learning
CATE vs ATE: who benefits, not only whether the average helps. Causal forests and metalearners with honest splitting; personalized treatment only under overlap, multiplicity control, and validation — without overclaiming precision medicine.
Open study →- ai-ml
- causal-ml
- cate
- causal-forests
- interview
- health
- AI / ML
Gradient Boosted Trees — XGBoost / LightGBM / CatBoost Tradeoffs
Cluster · Choosing ML Algorithms
“Just use XGBoost” is not a senior answer. Explain sequential residual fitting, the learning-rate × n_estimators tradeoff, early stopping, and honest differences among XGBoost, LightGBM, and CatBoost — then defend when GBM beats RF, and when calibration or latency make RF or linear wiser.
Open study →- ai-ml
- machine-learning
- ml
- interview
- xgboost
- lightgbm
- catboost
- gradient-boosting
- tabular
- AI / ML
Double Machine Learning — Nuisance Models, Orthogonalization & Cross-Fitting
Cluster · Causal ML & Double Machine Learning
Chernozhukov DML: residual-on-residual / orthogonal scores, flexible ML nuisances, and K-fold cross-fitting so a low-dimensional ATE stays valid. DML does not invent identification — it beats naive ML-on-treatment when the backdoor set is already right.
Open study →- ai-ml
- causal-ml
- double-ml
- orthogonalization
- cross-fitting
- interview
- AI / ML
Decision Trees — Splits, Interpretability & When They Fail
Cluster · Choosing ML Algorithms
A tree greedily partitions feature space into axis-aligned regions. That makes it readable — and brittle. Defend when a single tree is enough (debug, rules, small data, compliance) versus when it underfits or overfits, and why ensembles exist.
Open study →- ai-ml
- machine-learning
- ml
- interview
- decision-trees
- cart
- gini
- interpretability
- overfitting
- AI / ML
Choosing ML Algorithms — Problem Shape to Model Family
Cluster · Choosing ML Algorithms
Senior interviewers do not want an algorithm encyclopedia. They want you to map problem shape to model family under constraints: tabular vs sequence, interpretability vs accuracy, latency, data size, and drift. Start with a baseline; ship the simplest model that meets the metric and the ops budget.
Open study →- ai-ml
- machine-learning
- ml
- interview
- algorithm-selection
- decision-trees
- random-forest
- xgboost
- time-series
- baselines
- AI / ML
Causal Time Series — ITS, Synthetic Control & Diff-in-Diff Over Time
Cluster · Causal ML & Double Machine Learning
Interrupted time series, synthetic control, and DiD over time estimate counterfactuals after a shock. Parallel trends and pre-fit beat a low forecast MAPE. Prediction under status quo is a different question — do not re-teach ARIMA, Prophet, or GBM.
Open study →- ai-ml
- causal-ml
- time-series
- synthetic-control
- did
- interview
- AI / ML
Causal ML in Health — Outcomes, Bias, Ethics & Validation
Cluster · Causal ML & Double Machine Learning
Health causal ML is a safety rail: endpoints, selection bias, confounding by indication, fairness, and validation before acting on an estimate. The bravest senior answer is often not deploying a CATE yet.
Open study →- ai-ml
- causal-ml
- health
- ethics
- bias
- validation
- interview
- AI / ML
Causal ML & Double Machine Learning — From Association to Effect
Cluster · Causal ML & Double Machine Learning
Hub: correlation is not causation. Defend ATE, ATT, and CATE first, then identification, then estimators. Predictive ML is the wrong tool for treatment decisions; this cluster maps graphs, DML, CATE forests, causal time series, and health ethics.
Open study →- ai-ml
- causal-ml
- double-ml
- interview
- ate
- cate
- health
- AI / ML
Causal Graphs & Identification — DAGs, Confounders & Backdoor
Cluster · Causal ML & Double Machine Learning
DAGs make identification inspectable before any DML fit. Distinguish confounders, colliders, and mediators; apply backdoor and positivity; name unmeasured confounding and confounding by indication.
Open study →- ai-ml
- causal-ml
- dags
- identification
- interview
- health
- AI / ML
Validation & Repair Loops for LLM Outputs
Cluster · Structured outputs
Prefer constrained decoding for syntax. Use a bounded validate→feedback→regenerate loop for business rules or legacy models. Never repair refusals; never execute tools on invalid JSON.
Open study →- ai-ml-systems
- llms
- json-schema
- AI / ML
Tool/Function Calling vs Structured Outputs
Cluster · Structured outputs
Tools execute side effects and fetch live data; structured outputs constrain a final JSON contract; agent loops mix both. Pick by side effects, latency, and security — not by habit.
Open study →- ai-ml-systems
- llms
- agents
- AI / ML
Streaming Structured Output
Cluster · Structured outputs
Streaming tokens of JSON improves UX but partial JSON is invalid. Incremental parsers paint completed fields; commit side effects only after final schema validation. Constrained decoding still applies per token.
Open study →- ai-ml-systems
- llms
- json-schema
- AI / ML
JSON Schema Strictness for Structured Outputs
Cluster · Structured outputs
Strict structured outputs need a closed schema: root object, every property required, additionalProperties false, provider-supported subset, schema_version, null for absence.
Open study →- ai-ml-systems
- llms
- json-schema
- AI / ML
Grammars & CFG Constrained Decoding
Cluster · Structured outputs
Compile a CFG/GBNF to an automaton, mask illegal next tokens, keep a parse stack. Same idea as JSON Schema SO, but you can constrain SQL, arithmetic, or custom DSLs.
Open study →- ai-ml-systems
- llms
- json-schema
- AI / ML
Structured Outputs / Constrained Decoding
Cluster · Structured outputs
Strict JSON Schema vs JSON mode; CFG token masking; required fields + additionalProperties:false; Pydantic/Zod; refusals/truncation still break validity.
Open study →- ai-ml-systems
- llms
- json-schema