WORLD 04 POST-TRAINING & DISTILLATION
Watch raw capability become useful behavior.
From base weights to a steerable, safe assistant across 8 modules: SFT, synthetic data, LoRA math, RLHF / DPO, verifiable-reward reasoning, safety, evals, and deployment.
JOURNEY · 8 MODULES · 24 SCENES
From a raw base model to a deployed, aligned system.
SFT Foundations
From raw base model to instruction follower: the SFT pipeline, loss masking, and the LoRA memory trade.
01 · SFT
Supervised Fine-Tuning Fundamentals
From base model to instruction-follower: the SFT pipeline that reshapes next-token distribution into assistant behavior.
OPEN EXHIBIT →02 · Loss
SFT Loss: Prompt Masking & Completion Targets
Cross-entropy scored only over completion positions — why prompt masking matters and what it does to effective epochs.
OPEN EXHIBIT →03 · PEFT
LoRA & QLoRA: Math and Memory Layouts
Rank-r adapters, trainable-parameter share, and the 4-bit base + bf16 LoRA memory bill vs. a full fine-tune.
OPEN EXHIBIT →Data Engineering & Synthetic Data
Curation heuristics, embedding clustering, synthetic generation loops, and the rejection gate.
04 · Curation
Data Curation & Quality Filtering
Heuristic filters, language ID, and taxonomy design for instruction, multi-turn, and safety data trees.
OPEN EXHIBIT →05 · Diversity
Embedding Clustering & Deduplication
MinHash dedup and embedding-space clustering that keeps topic coverage balanced across the post-training corpus.
OPEN EXHIBIT →06 · SDG
Self-Instruct, Evol-Instruct & Rejection Sampling
Synthetic data generation loops, model-collapse risk, and the quality gate that rejects low-fidelity generations.
OPEN EXHIBIT →SFT Deep Dive
Low-rank adaptation math, QLoRA memory layouts, catastrophic forgetting, and domain adaptation.
07 · LoRA
LoRA: Why Low-Rank Adaptation Works
The ΔW = BA decomposition, target-module selection, and parameter-count math for rank and target fraction.
OPEN EXHIBIT →08 · QLoRA
QLoRA Memory: 4-bit Base, Full-Pipeline Bills
NF4 quantization constants, PagedAdamW, and the VRAM ledger that fits a 7B tune on a single consumer GPU.
OPEN EXHIBIT →09 · CPT
Catastrophic Forgetting & Domain Adaptation
Rehearsal buffers, distillation bridges, and the retention math that keeps general capability while specializing.
OPEN EXHIBIT →Preference & Alignment
Bradley-Terry, RLHF with PPO, the DPO revolution, and the KTO / IPO / SimPO frontier.
10 · BT
The Bradley-Terry Preference Model
P(y_w ≻ y_l) = σ(r_w − r_l): the probability model underpinning RLHF and every preference-optimization method.
OPEN EXHIBIT →11 · RLHF
RLHF: Reward Models & PPO with KL Anchors
Training a reward model from pairs, then PPO's clipped objective, GAE, and the KL penalty tethering π to π_ref.
OPEN EXHIBIT →12 · DPO
DPO and Its Heirs: KTO, IPO, SimPO
Eliminating the reward model, the implicit-reward insight, and the variant family that followed.
OPEN EXHIBIT →Reasoning & Verifiable Rewards
RLVR, GRPO without a critic, chain-of-thought elicitation, and pass@k gains on math and code.
13 · RLVR
RLVR: Rewards You Can Verify
Math checkers, compilers, and unit tests as deterministic reward oracles — no human labels in the loop.
OPEN EXHIBIT →14 · GRPO
GRPO: Group-Relative Advantages, No Critic
Sampling a group, normalizing rewards within it, and deleting the value network from the memory bill.
OPEN EXHIBIT →15 · CoT
Chain-of-Thought & Test-Time Compute
System 1 vs. System 2, self-consistency majority voting, and pass@k scaling on MATH and HumanEval.
OPEN EXHIBIT →Safety & Constitutional AI
Red-teaming loops, refusal calibration, RLAIF critics, and the safety/helpfulness frontier.
16 · Red-Team
Red-Teaming & Jailbreak Decay
Automated adversarial loops that attack the model, and the exponential decay of jailbreak success per round.
OPEN EXHIBIT →17 · Refusal
Refusal Calibration & the Over-Refusal Trap
Training refusals without breaking helpfulness: the quadratic cost of over-refusal on the frontier curve.
OPEN EXHIBIT →18 · CAI
Constitutional AI & AI Critics (RLAIF)
Self-critique, revision loops, and judge agreement with human panels as the violation rate decays.
OPEN EXHIBIT →Evaluation & Benchmarking
Contamination guardrails, LLM-as-a-judge biases, and honest model reporting.
19 · Contam
Benchmark Contamination & Guardrails
Seen vs. unseen accuracy uplift as the contamination tell, and the guardrails that keep evals honest.
OPEN EXHIBIT →20 · Judge
LLM-as-a-Judge: Bias Audit
Positional flips, length bias, and judge Elo deltas — measuring and mitigating the referee's own errors.
OPEN EXHIBIT →21 · Stats
Reporting Uncertainty: CIs & Elo
Confidence-interval half-widths, win-rate to Elo conversion, and why single-point leaderboard scores mislead.
OPEN EXHIBIT →Advanced Capabilities & Deployment
Multimodal alignment, agentic tool use, post-training quantization, and deployment prep.
22 · Agents
Multimodal & Agentic Alignment
Interleaved image/audio token mix and the composite tool-call success rate of syntax plus argument fidelity.
OPEN EXHIBIT →23 · Quant
Post-Training Quantization: GPTQ, AWQ, FP8
r-bit grouped weight error, quality retained, and the VRAM you actually need at bits-per-weight.
OPEN EXHIBIT →24 · Ship
Deployment: Distillation & Serving Prep
Teacher-student KL distillation, FP8 calibration curves, and the pipeline from checkpoint to serving.
OPEN EXHIBIT →