AI ENGVisual Encyclopedia

WORLD 04  POST-TRAINING & DISTILLATION

Watch raw capability become useful behavior.

From base weights to a steerable, safe assistant across 8 modules: SFT, synthetic data, LoRA math, RLHF / DPO, verifiable-reward reasoning, safety, evals, and deployment.

JOURNEY · 8 MODULES · 24 SCENES

From a raw base model to a deployed, aligned system.

MODULE 01

SFT Foundations

From raw base model to instruction follower: the SFT pipeline, loss masking, and the LoRA memory trade.

MODULE 02

Data Engineering & Synthetic Data

Curation heuristics, embedding clustering, synthetic generation loops, and the rejection gate.

MODULE 03

SFT Deep Dive

Low-rank adaptation math, QLoRA memory layouts, catastrophic forgetting, and domain adaptation.

MODULE 04

Preference & Alignment

Bradley-Terry, RLHF with PPO, the DPO revolution, and the KTO / IPO / SimPO frontier.

MODULE 05

Reasoning & Verifiable Rewards

RLVR, GRPO without a critic, chain-of-thought elicitation, and pass@k gains on math and code.

MODULE 06

Safety & Constitutional AI

Red-teaming loops, refusal calibration, RLAIF critics, and the safety/helpfulness frontier.

MODULE 07

Evaluation & Benchmarking

Contamination guardrails, LLM-as-a-judge biases, and honest model reporting.

MODULE 08

Advanced Capabilities & Deployment

Multimodal alignment, agentic tool use, post-training quantization, and deployment prep.