The K3 thesis
Separate pretraining scale, test-time effort, active capacity, context, multimodality, and evidence levels before meeting the mechanisms.
Learn through computational animation, rigorous derivations, executable miniatures, retained results, and explicit evidence boundaries.
Follow the paper from thesis to bounded conclusion.
Separate pretraining scale, test-time effort, active capacity, context, multimodality, and evidence levels before meeting the mechanisms.
Follow one token through KDA, Gated MLA, AttnRes, Stable LatentMoE, native vision, and Per-Head Muon.
Connect data curation, scaling-law allocation, numerical controls, and progressive extension to one million tokens.
Trace SFT, effort-conditioned RL, multi-teacher on-policy distillation, environments, task synthesis, QAT, and speculative decoding.
See how kernels, distributed scans, expert placement, memory lifetimes, long-context environments, hybrid caches, and fleet scheduling make scale operable.
Read capability profiles without inventing an overall score, inspect selected cases as trajectories, and stop at the strongest supportable claim.