Learn and practice the building blocks of Kimi K3
Original paper ↗
Paper-aligned learning system

Understand the system, not merely the diagram.

Learn through computational animation, rigorous derivations, executable miniatures, retained results, and explicit evidence boundaries.

The learning route

Follow the paper from thesis to bounded conclusion.

Course 0 · Section 1

The K3 thesis

Separate pretraining scale, test-time effort, active capacity, context, multimodality, and evidence levels before meeting the mechanisms.

Cinematic introduction + evidence contract
Course 1 · Section 2

Model architecture

Follow one token through KDA, Gated MLA, AttnRes, Stable LatentMoE, native vision, and Per-Head Muon.

Cinematic course + five executed miniatures
Course 2 · Section 3

Pretraining

Connect data curation, scaling-law allocation, numerical controls, and progressive extension to one million tokens.

Complete course surface
Course 3 · Section 4

Post-training and agentic RL

Trace SFT, effort-conditioned RL, multi-teacher on-policy distillation, environments, task synthesis, QAT, and speculative decoding.

Complete course surface
Course 4 · Section 5

Frontier infrastructure

See how kernels, distributed scans, expert placement, memory lifetimes, long-context environments, hybrid caches, and fleet scheduling make scale operable.

Paper-complete course + systems prerequisite bridge
Courses 5A–5B · Sections 6–8

Evaluation to conclusion

Read capability profiles without inventing an overall score, inspect selected cases as trajectories, and stop at the strongest supportable claim.

Evaluation and trajectory synthesis complete

Evidence labels are part of the lesson

Executed hereA deterministic miniature or analysis ran in this repository.
Animated interpretationMotion explains a mechanism; it is not an experiment or benchmark.
Reported by paperThe full-scale claim belongs to Kimi K3 and is not independently reproduced here.