LLM Post-Training
Studies in this cluster, in series order. Each one keeps its own URL.
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
LLM Post-Training
6 studies- 1.LLM Post-Training - From Base Model to Assistant: SFT, Preference Tuning, PEFT, Distillation & EvalsHub: base model -> assistant pipeline (SFT, RLHF/PPO, DPO-family and GRPO, LoRA/QLoRA, distillation, eval gate); techniques compared by signal needed, cost and failure mode; runnable pipeline toy + technique picker; decision chart.
- 2.Supervised Fine-Tuning (SFT) - Instruction Data, Chat Templates, Packing & Loss MaskingSFT data formats, chat templates and special tokens, assistant-only loss masking, packing with attention boundaries, truncation; data quality over volume; runnable masking + packing demos; template/masking bug failure paths.
- 3.Reward Models & RLHF with PPO - Bradley-Terry, KL Penalty, Critics & Reward HackingBradley-Terry reward models from pairwise preferences, PPO with KL penalty to the SFT reference, critic/GAE, four-models-in-memory cost, reward hacking and over-optimisation; runnable RM hacking demo + PPO step.
- 4.Preference Optimization - DPO, IPO, ORPO, SimPO, KTO & GRPO Compared with PPODPO derivation from the RLHF objective; IPO, KTO, ORPO, SimPO variants (reference-free, unpaired data); GRPO group-relative advantages with verifiable rewards; all compared with PPO by signal, cost, failure mode; runnable losses + GRPO advantages.
- 5.LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter MergingLoRA math (W + alpha/r BA), rank/alpha/target modules, rsLoRA/DoRA notes, QLoRA NF4 + double quantization + paged optimizers, merging vs keeping adapters, when PEFT is the wrong choice; runnable LoRA/NF4 demo + memory budget calculator.
- 6.Distillation & LLM Evals - Teacher-Student Training, Benchmarks, LLM-as-Judge, Regression Gates & ContaminationSequence-level vs logit-level vs on-policy distillation and reasoning-trace distillation; benchmarks, LLM-as-judge biases, regression gates with per-slice thresholds and CIs, contamination checks; runnable distill/contamination + regression gate.