Animated Kimi K3 architecture diagram
A token moves through Kimi K3's multimodal input, hybrid attention backbone, depth routing, expert routing, and output head.
ONE TOKEN · xₜSEQUENCE → DEPTH → WIDTH
A CONTINUOUS REPRESENTATION, RE-ROUTED AS IT GROWS
token journey
next-token distribution
t
ONE TOKEN · THREE ROUTING PROBLEMS
xₜ
one hidden state
SEQUENCE 1M causal positions
SEQUENCE memory & global read
DEPTH 93 transformations
DEPTH retrieve earlier layers
WIDTH 896 expert choices
WIDTH sparse capacity
FIRST, PUT EVERY MODALITY ON ONE CAUSAL TIMELINE
TEXT token IDs … the moon is
IMAGE / VIDEO
patches
MOONVIT-V2 visual encoder 27-layer ViT 2×2 shuffle → MLP
MLP projector visual → d shared width
ONE STREAM x₁ … xₜ … xₜ₊₁
causal sequence shared backbone
BUILD THE RHYTHM, THEN COMPRESS ITS REPEAT
23 repeating groups three recurrent KDA layers, then one global MLA aperture
KDA ① carry + correct
KDA ② carry + correct
KDA ③ carry + correct
GATED MLA ④ open global aperture
t
RECURRENT STATE
FULL-PREFIX READ
THE PATTERN, COMPILED
KDA
KDA
KDA
MLA
× 23
MLA
69 KDA + 24 MLA = 93 attention layers
Every attention layer is followed by Stable LatentMoE.
LOCAL MEMORY: A FIXED MATRIX THAT PREDICTS, FORGETS, AND CORRECTS
KDA state, per head
Sₜ ∈ ℝᵈᵏ ˣ ᵈᵥ · FIXED SHAPE
fixed-size state
recurrent memory carry S across chunks
ordinary softmax attention
GROWING KV CACHE
… T
cache grows with T one K,V pair per token
1 · PROBE kₜ reads S̄ₜ
2 · CORRECT eₜ = vₜ − S̄ₜᵀkₜ
3 · DECAY Diag(αₜ) Sₜ₋₁
4 · WRITE βₜ kₜ eₜᵀ → Sₜ
NOTEBOOK TRACE · THE ANIMATION BECOMES A MEASUREMENT
final S
αₜ · channel 0
miniature output [2,16,64] · sequential ≈ chunkwise · state shape stays fixed
COMPRESS WHAT PERSISTS · EXPAND ONLY WHEN CONTENT IS READ
ONE TOKEN ENTERS · FOUR COMPUTATIONAL MOVES
1 · FULL-WIDTH STATE
xᵢ ∈ ℝ⁶⁴
model width 64
Wc
2 · CACHE LATENT
cᵢ ∈ ℝ¹⁶
only this persists
3 · GLOBAL LOOKUP
Qₜ asks
softmax
over all cᵢ
content, not distance
4 · RECONSTRUCT + GATE
cᵢ → Kᵢ, Vᵢ
MLA(xₜ)
× σ(Wg xₜ) · output gate
return to width 64
EXECUTED NOTEBOOK ECHO · SEE THE SCALE CHANGE, THEN CHECK THE SHAPE
64
16
64
output shape [2, 16, 64]
KDA carries · MLA retrieves
shape verified · motion interpreted
IN DEPTH, REPLACE A SINGLE PIPE WITH A SELECTIVE RETRIEVAL
ordinary residual
hₗ = hₗ₋₁ + fₗ(hₗ₋₁)
embedding
layer 01
layer 02
layer 03
… deeper
one running mixture history loses its labels
BLOCK ATTNRES · DEPTH SOFTMAX
8 compact block summaries
b₀ embedding
b₁
b₂
b₃
… b₈
LEARNED wₗ pseudo-query scores RMSNorm(bᵢ)
softmax weights
Σ αᵢbᵢ
The layer receives a targeted blend.
Paper: O(8d) summaries, not 93 full outputs.
Current partial block is also addressable.
NOTEBOOK · TWO TOKENS, THREE DEPTH SOURCES
t₀ [.665, .245, .090]
t₁ [.090, .665, .245]
Σ = 1
limit · supplied sources, not paper blocks
IN WIDTH, SCORE MANY SPECIALISTS — THEN EXECUTE ONLY SIXTEEN
TOKEN xₜ two paths begin
always-on shared path
routed latent path
shared path stays experts specialize
ROUTER SCORE SLICE · 896 TOTAL EXPERTS
sⱼ + bⱼ ranks experts for this token
τ threshold
shown scores … 880 more
SELECTED LATENT EXPERTS Top 16 of 896
RMSNorm → W↑ + shared QB balances selection; bias freezes for inference.
paper · shared full-width path + sparse update
LOCAL MINIATURE · PARTIAL MECHANISM
local |SiTU| 15.997 ≤ 16
top-2 [0,1] · [2,1]
EMA bias [0,−.001,0,.001]
no local shared path
TRAINING LENS · ORTHOGONALIZE EACH ATTENTION HEAD WITHOUT MIXING ITS NEIGHBORS
A FUSED QKV GRADIENT IS STORAGE — NOT THE UNIT OF GEOMETRY
1 · PACKED GRADIENT
G ∈ ℝ²⁴×⁸
Q K V
one fused tensor
reshape
2 · RESTORE HEAD BOUNDARIES
3 × 2 heads × 4×8
Q₀
Q₁
K₀
K₁
V₀
V₁
six independent matrices
3 · FIVE NEWTON–SCHULZ STEPS
Xₖ
→ Xₖ₊₁
matrix polynomial
head-local update
4 · REPACK UPDATE
ΔWₕᵤᵒⁿ
repack same shape
geometry preserved
NOTEBOOK · DETERMINISTIC 2-HEAD FUSED QKV
head-local error 0.7348
whole-slice 1.1796
not a convergence claim
SYNTHESIS: ROUTED INFORMATION BECOMES THE NEXT-TOKEN DISTRIBUTION
93 LAYERS h_final sequence + global depth + experts
FINAL APERTURE Gated MLA final global read then RMSNorm
LM HEAD Wout 160K logits softmax
NEXT-TOKEN DISTRIBUTION
probability sharpens the model emits one token, then the journey repeats.
next token
sequence carries evidence · depth retrieves the right abstraction · width specializes the transform
Motion has meaning. The glowing token is the same representation as it moves through each route.