Concept and prerequisite library
Use this library when a course names an unfamiliar mechanism. Each entry separates the general idea from K3's adaptation and preserves its evidence boundary.
Agents
Autonomous agents
Intuition. Decompose and execute objectives without a supplied procedure.
Definition. A policy plans acts observes recovers and terminates under constraints.
Formula or process. objective -> plan -> actions -> verified state
Learn first. tool-use
K3 adaptation. Trained with AET.
Common misconception. Autonomy is not absence of constraints.
Boundary. Long horizons compound error.
Sources. Kimi K3: Open Frontier Intelligence
Long-horizon agents
Intuition. Maintain progress across many dependent actions.
Definition. Agents solve tasks whose consequences and required evidence span extended trajectories.
Formula or process. tau=(s0,a0,...,sT)
Learn first. agentic-environments
K3 adaptation. Partial rollout preserves long trajectories.
Common misconception. More steps do not imply better reasoning.
Boundary. Cost staleness and recovery dominate.
Sources. Kimi K3: Open Frontier Intelligence
Persistent state
Intuition. Let actions affect what later steps observe.
Definition. An environment retains world changes across turns events and simulated time.
Formula or process. s_(t+1)=f(s_t,a_t,e_t)
Learn first. agentic-environments
K3 adaptation. Powers multi-application assistant worlds.
Common misconception. Context text is not the only state.
Boundary. Reset and reproducibility are hard.
Sources. Kimi K3: Open Frontier Intelligence
Tool use
Intuition. Act through structured external interfaces.
Definition. A policy emits arguments to an API and conditions on returned observations.
Formula or process. a=tool(args); observe(result)
Learn first. None
K3 adaptation. Central to agent trajectories.
Common misconception. Calling a tool does not verify success.
Boundary. Schemas and failures vary.
Sources. Kimi K3: Open Frontier Intelligence
Web development
Intuition. Produce an executable interactive artifact from a specification.
Definition. Implement build run and verify frontend or full-stack software in a sandbox.
Formula or process. specification -> code -> running artifact
Learn first. tool-use
K3 adaptation. Forms a diverse RL task family.
Common misconception. Visual resemblance alone does not prove functionality.
Boundary. Framework and evaluator coverage vary.
Sources. Kimi K3: Open Frontier Intelligence
Algorithms
Associative prefix scan
Intuition. Compute every prefix using a tree because grouping does not change the answer.
Definition. For associative operator circle scan returns x1 and x1 circle x2 through the full prefix.
Formula or process. p_i = x_1 circle ... circle x_i
Learn first. None
K3 adaptation. KCP scans affine segment summaries.
Common misconception. The operator need not be commutative.
Boundary. Summary size and communication still matter.
Sources. Kimi K3: Open Frontier Intelligence
Online softmax
Intuition. Normalize a stream without storing every unnormalized exponential.
Definition. Maintain a running maximum and rescaled exponential sum so new score blocks can be merged exactly.
Formula or process. m'=max(m,x); l'=exp(m-m')l+sum exp(x-m')
Learn first. attention-mechanism
K3 adaptation. Block AttnRes can stream depth-score normalization with bounded working memory.
Common misconception. Online changes scheduling not the mathematical softmax result.
Boundary. Source representations still need a storage or recomputation policy.
Sources. Kimi K3: Open Frontier Intelligence
Architecture
Attention residuals
Intuition. Retrieve earlier depth representations using learned attention weights.
Definition. Depth summaries are scored normalized and mixed instead of only added sequentially.
Formula or process. h_l = sum_i alpha_i b_i
Learn first. attnres
K3 adaptation. Block AttnRes requires specialized serving gather kernels.
Common misconception. It is depth attention rather than token attention.
Boundary. Summary construction affects fidelity.
Sources. Kimi K3: Open Frontier Intelligence
Latent experts
Intuition. Execute routed expert transformations in a compressed representation space.
Definition. Down-project tokens route through sparse experts then up-project the combined latent result.
Formula or process. x -> z -> sparse experts -> x space
Learn first. mixture-of-experts, low-rank-compression
K3 adaptation. Stable LatentMoE reduces active expert compute width.
Common misconception. Latent routing does not eliminate expert communication.
Boundary. Compression can bottleneck capacity.
Sources. Kimi K3: Open Frontier Intelligence, LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
Mixture of experts
Intuition. Activate a small parameter subset for each token.
Definition. A router selects experts whose outputs are weighted and combined.
Formula or process. y = sum_i in top-k p_i E_i(x)
Learn first. moe-routing
K3 adaptation. Stable LatentMoE routes sixteen of 896 reported experts.
Common misconception. Sparse FLOPs do not imply sparse communication.
Boundary. Routing skew harms utilization.
Sources. Kimi K3: Open Frontier Intelligence
Multi-token prediction
Intuition. Predict multiple future tokens from one position.
Definition. Auxiliary prediction layers learn representations useful beyond the immediate next token.
Formula or process. predict x_(t+1..t+k)
Learn first. None
K3 adaptation. The MTP layer initializes the draft.
Common misconception. MTP alone is not speculative decoding.
Boundary. Future targets become progressively uncertain.
Sources. Kimi K3: Open Frontier Intelligence
Multimodal encoders
Intuition. Convert non-text inputs into representations a shared backbone can consume.
Definition. Modality-specific networks extract and project features before shared sequence processing.
Formula or process. image -> features -> projected tokens
Learn first. vision-transformer
K3 adaptation. MoonViT-V2 work is decomposed around the LM pipeline.
Common misconception. Projection does not erase encoder cost.
Boundary. Variable input size creates load imbalance.
Sources. Kimi K3: Open Frontier Intelligence
Vision transformers
Intuition. Treat image patches as a sequence for transformer processing.
Definition. Images become patch embeddings processed by attention and feed-forward layers.
Formula or process. image -> patches -> token sequence
Learn first. vision-transformer
K3 adaptation. Dynamic context groups partition large MoonViT workloads.
Common misconception. Patch count is not constant across resolutions.
Boundary. Attention and load grow with visual tokens.
Sources. Kimi K3: Open Frontier Intelligence
Attention Memory
Attention complexity
Intuition. Global pairwise interaction becomes expensive as sequence length grows.
Definition. Standard full attention forms token-pair scores, while recurrent alternatives maintain fixed-size state.
Formula or process. full attention compute O(T^2); recurrent state update O(T)
Learn first. softmax-attention
K3 adaptation. K3 combines recurrent KDA groups with periodic global MLA layers.
Common misconception. Linear recurrent compute does not guarantee unlimited memory quality.
Boundary. Wall-clock behavior also depends on kernels, parallelism, and memory traffic.
Sources. Kimi K3: Open Frontier Intelligence
Attention mechanism
Intuition. Retrieve a weighted mixture of values using query-key compatibility.
Definition. Attention scores candidate keys against a query normalizes the scores and combines their values.
Formula or process. Attention(Q,K,V)=softmax(QK^T/sqrt(d))V
Learn first. None
K3 adaptation. K3 applies attention across sequence positions and separately across depth.
Common misconception. Attention is a retrieval rule not a guarantee of reasoning.
Boundary. Cost and memory depend on the attention axis and implementation.
Sources. Attention Is All You Need
Context window
Intuition. Bound how many tokens can jointly participate in one model invocation.
Definition. The context window is the maximum admitted sequence length, distinct from proven useful recall distance.
Formula or process. T <= T_max
Learn first. None
K3 adaptation. K3 progresses through 8K, 64K, 256K, and 1M stages.
Common misconception. Capacity does not imply effective distant reasoning.
Boundary. Useful context depends on training evidence and architecture.
Sources. Kimi K3: Open Frontier Intelligence
DeepSeek Sparse Attention
Intuition. Reduce long-context attention work by selecting a smaller token subset for each query.
Definition. A learned or structured indexer identifies relevant positions before computing a sparse attention result.
Formula or process. query -> select positions S -> attention over S
Learn first. softmax-attention, attention-complexity
K3 adaptation. K3 discusses DSA as a comparison point while choosing hybrid KDA-MLA execution.
Common misconception. DSA selects token interactions whereas KDA compresses history into recurrent state.
Boundary. This course does not reproduce DSA's kernel or quality results.
Sources. Kimi K3: Open Frontier Intelligence, DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Delta rule
Intuition. Write only what the current memory predicts incorrectly.
Definition. A fast-weight update subtracts the value predicted at a key and writes the residual error back at that key.
Formula or process. S_t=S_{t-1}+beta k(v-S_{t-1}^T k)^T
Learn first. linear-attention, fast-weights
K3 adaptation. KDA combines delta correction with lower-bounded channel-wise decay.
Common misconception. It is not ordinary gradient descent through model weights.
Boundary. Fixed state can overwrite or compress history.
Sources. Parallelizing Linear Transformers with the Delta Rule over Sequence Length
Fast weights and DeltaNet
Intuition. Write token associations into a matrix that acts as temporary memory.
Definition. A recurrent state is updated by outer products; the delta rule writes prediction error instead of blindly accumulating values.
Formula or process. S_t = S_{t-1} + beta_t k_t(v_t - S_{t-1}^T k_t)^T
Learn first. softmax-attention
K3 adaptation. KDA adds channel-wise lower-bounded decay and systems-oriented chunkwise execution.
Common misconception. Fixed state size does not guarantee indefinite semantic retention.
Boundary. Repository miniatures do not benchmark million-token memory quality.
Sources. Parallelizing Linear Transformers with the Delta Rule over Sequence Length
Global causal attention
Intuition. Let each token address any permitted earlier token directly.
Definition. A causal attention row scores all prefix positions before normalizing and mixing their values.
Formula or process. y_t=sum_{i<=t} softmax(q_t k_i) v_i
Learn first. attention-mechanism
K3 adaptation. Periodic Gated MLA layers provide global prefix lookup.
Common misconception. Global access does not imply equal use of all tokens.
Boundary. Standard cache storage grows with prefix length.
Sources. Attention Is All You Need
Kimi Delta Attention
Intuition. Carry a fixed correctable memory rather than a growing list of token KV pairs.
Definition. KDA decays state rows, measures prediction error, writes a rank-one correction, and reads with the query.
Formula or process. Sbar_t=Diag(alpha_t)S_{t-1}; e_t=v_t-Sbar_t^T k_t; S_t=Sbar_t+beta_t k_t e_t^T
Learn first. fast-weights
K3 adaptation. Three KDA layers alternate with each Gated MLA layer in the hybrid backbone.
Common misconception. The decay lower bound is a numerical parameterization, not proof of 1M-token semantic stability.
Boundary. The reference implementation is sequential inside chunks, not the production FlashKDA kernel.
Sources. Kimi K3: Open Frontier Intelligence
KV cache
Intuition. Save prior keys and values so decoding does not recompute them.
Definition. Autoregressive inference retains projected keys and values for all reusable prior tokens.
Formula or process. storage proportional to sequence length × layers × KV width
Learn first. softmax-attention
K3 adaptation. K3 combines an MLA latent KV cache with fixed KDA recurrent states.
Common misconception. KV storage is linear in context length; ordinary full-attention prefill compute is quadratic.
Boundary. Actual memory depends on precision, head sharing, paging, and compression.
Sources. Attention Is All You Need
Latent attention
Intuition. Cache compact token coordinates and reconstruct attention projections on demand.
Definition. Low-rank latent representations replace full per-head key and value storage.
Formula or process. c_t = W_c x_t
Learn first. low-rank-compression
K3 adaptation. Gated MLA supplies the KV half of the hybrid serving cache.
Common misconception. Latent cache and recurrent state have different granularities.
Boundary. Reconstruction still costs compute.
Sources. Kimi K3: Open Frontier Intelligence
Linear attention
Intuition. Carry a fixed-size summary instead of materializing every query-key pair.
Definition. A feature-map or recurrent formulation reassociates attention into state updates with sequence-linear computation.
Formula or process. S_t=S_{t-1}+phi(k_t)v_t^T; y_t=phi(q_t)^T S_t
Learn first. attention-mechanism
K3 adaptation. KDA adds delta-rule correction and channel-wise decay.
Common misconception. Linear-time state is not lossless full attention.
Boundary. Expressivity depends on state size update rule and feature map.
Sources. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Multi-head attention
Intuition. Run several attention subspaces in parallel.
Definition. Learned projections partition queries keys and values into heads whose outputs are concatenated and projected.
Formula or process. MHA=Concat(head_1...head_H)W_o
Learn first. attention-mechanism
K3 adaptation. Gated MLA reconstructs head-specific views from compact latents.
Common misconception. Heads are learned subspaces not guaranteed human-interpretable roles.
Boundary. Fused storage can hide head-local geometry.
Sources. Attention Is All You Need
Positional encoding
Intuition. Provide or induce information about token order and distance.
Definition. Explicit embeddings, rotations, biases, or recurrence let sequence models distinguish positions.
Formula or process. representation_t = content_t + or compose position_t
Learn first. None
K3 adaptation. K3 uses NoPE and therefore avoids positional rescaling during extension.
Common misconception. NoPE does not mean order is unknowable to the network.
Boundary. Extrapolation behavior depends on the full architecture and curriculum.
Sources. Kimi K3: Open Frontier Intelligence
Recurrent state
Intuition. Summarize prior sequence information in a fixed-size object updated over time.
Definition. A transition maps incoming state and current input to outgoing state and output.
Formula or process. S_t = F(S_t-1,x_t)
Learn first. fast-weights
K3 adaptation. KDA checkpoints represent causal prefix boundaries.
Common misconception. Fixed size does not imply lossless recall.
Boundary. Checkpoints can be large even when sequence-independent.
Sources. Kimi K3: Open Frontier Intelligence
Softmax attention
Intuition. A token scores earlier tokens and blends their values.
Definition. Query-key similarities are normalized with softmax and used as value weights.
Formula or process. Attention(Q,K,V) = softmax(QK^T / sqrt(d_k))V
Learn first. None
K3 adaptation. Periodic Gated MLA preserves global softmax interaction inside a mostly recurrent backbone.
Common misconception. Attention compute and KV-cache storage have different scaling laws.
Boundary. This definition alone does not describe efficient kernels or cache compression.
Sources. Attention Is All You Need
Data
Instruction data
Intuition. Demonstrate desired responses and action sequences.
Definition. Curated prompt-response or trajectory records supervise policy behavior.
Formula or process. (instruction,demonstration)
Learn first. None
K3 adaptation. Uses verified XTML trajectories.
Common misconception. Quality matters beyond volume.
Boundary. Coverage restricts cold-start behavior.
Sources. Kimi K3: Open Frontier Intelligence
Knowledge graphs
Intuition. Organize concepts and their directed relationships.
Definition. Nodes represent concepts and edges encode typed or hierarchical relations.
Formula or process. G=(V,E)
Learn first. None
K3 adaptation. A self-evolving DAG guides retrieval.
Common misconception. A graph is not automatically a source of truth.
Boundary. Node quality and duplication matter.
Sources. Kimi K3: Open Frontier Intelligence
Depth Routing
Block Attention Residuals
Intuition. Let each layer retrieve the useful abstraction from earlier depth instead of accepting one accumulated mixture.
Definition. A learned pseudo-query produces softmax weights over embeddings and preceding block summaries.
Formula or process. alpha_{i->l}=softmax(w_l·RMSNorm(b_i)/sqrt(d)); h_l=sum_i alpha_{i->l}b_i
Learn first. residual-connections, rmsnorm
K3 adaptation. K3 uses eight compact block summaries to control memory overhead.
Common misconception. AttnRes operates across network depth, not sequence positions.
Boundary. Miniature attention weights do not reproduce learned full-model routing patterns.
Sources. Kimi K3: Open Frontier Intelligence
Residual connections
Intuition. Carry a representation around a transformation so optimization and information flow remain tractable.
Definition. A block adds its transformed input to an existing residual stream.
Formula or process. h_l = h_{l-1} + f_l(h_{l-1})
Learn first. None
K3 adaptation. AttnRes replaces blind accumulation with learned retrieval across depth.
Common misconception. An additive stream does not preserve individually addressable earlier representations.
Boundary. Residual behavior depends on normalization and block design.
Sources. Deep Residual Learning for Image Recognition
Evaluation
Agent harnesses
Intuition. The scaffold can amplify or constrain the model.
Definition. A harness supplies prompts tools memory context management and an action loop around a model.
Formula or process. trajectory = harness(model, task, tools)
Learn first. agentic-environments
K3 adaptation. Coding baselines run through Kimi Code Claude Code or Codex depending on task and model.
Common misconception. An agent score measures only the base model.
Boundary. Cross-harness comparisons can confound model and scaffold quality.
Sources. Kimi K3: Open Frontier Intelligence
Benchmark design
Intuition. Choose tasks that expose the failure mode you care about.
Definition. Benchmark design specifies a target construct task distribution interface scoring rule and contamination controls.
Formula or process. benchmark = tasks + interface + metric + controls
Learn first. capability-evaluation
K3 adaptation. Internal K3 suites are refreshed to follow evolving coding agent and conversational failures.
Common misconception. Harder questions automatically make a better benchmark.
Boundary. Internal suites are harder for outsiders to audit.
Sources. Kimi K3: Open Frontier Intelligence
Benchmark interpretation
Intuition. Read a number together with its metric and task distribution.
Definition. Interpretation connects a reported score to what was sampled how it was scored and which comparisons are valid.
Formula or process. claim <= evidence(protocol, sample, metric)
Learn first. comparative-evaluation
K3 adaptation. Course 5A separates measured gaps from causal hypotheses.
Common misconception. A lead explains why the lead occurred.
Boundary. Aggregate scores conceal item-level failure modes.
Sources. Kimi K3: Open Frontier Intelligence
Benchmarking
Intuition. Compare implementations under controlled conditions.
Definition. Repeat a defined workload and summarize correctness and performance against baselines.
Formula or process. score(candidate,reference)
Learn first. None
K3 adaptation. Rewards kernels against experts and roofline.
Common misconception. Benchmarks can be gamed.
Boundary. Results may not transfer workloads.
Sources. Kimi K3: Open Frontier Intelligence
Capability evaluation
Intuition. Ask what kind of work a system can reliably complete.
Definition. Capability evaluation samples tasks that require a target skill and measures outcomes under a specified interface.
Formula or process. capability ~= performance(task distribution, protocol)
Learn first. evaluation-protocols
K3 adaptation. Public and internal suites probe reasoning coding agency vision and conversation.
Common misconception. One benchmark does not define a general capability.
Boundary. Coverage is always incomplete and distribution-dependent.
Sources. Kimi K3: Open Frontier Intelligence
Case-study method
Intuition. Inspect a complete task trajectory to understand capability and failure recovery.
Definition. A bounded case study records objective environment actions artifacts verification and limitations for one situated task.
Formula or process. claim <= observed trajectory + verification
Learn first. evaluation-limitations
K3 adaptation. Section 7 demonstrates kernels compilers chips science knowledge work and media workflows.
Common misconception. One successful case estimates neither prevalence nor average reliability.
Boundary. Selection and reporting can hide failures.
Sources. Kimi K3: Open Frontier Intelligence
Comparative evaluation
Intuition. Compare like with like and preserve the denominator.
Definition. Comparative evaluation contrasts systems on the same task and metric while documenting configuration differences.
Formula or process. Delta_i = s_i(candidate) - s_i(reference)
Learn first. baselines, evaluation-protocols
K3 adaptation. K3 is read as a capability profile rather than one averaged rank.
Common misconception. Percent F1 Elo and pass@k are not interchangeable.
Boundary. Missing entries and heterogeneous metrics prevent a universal average.
Sources. Kimi K3: Open Frontier Intelligence
Cost efficiency frontier
Intuition. Prefer systems for which no alternative is both better and cheaper.
Definition. A point is Pareto-efficient when no comparator has at least its score at no greater cost with one strict improvement.
Formula or process. not exists j: score_j >= score_i and cost_j <= cost_i
Learn first. comparative-evaluation, inference-cost
K3 adaptation. K3 is reported on or near the frontier across four suites.
Common misconception. The cheapest model is automatically the most efficient.
Boundary. The frontier changes with pricing workload and effort setting.
Sources. Kimi K3: Open Frontier Intelligence
Cybersecurity evaluation
Intuition. Measure capability by operational difficulty and misuse relevance.
Definition. Cyber evaluations grade tasks such as vulnerability discovery and complete exploit development under controlled targets.
Formula or process. tier 1 discovery -> tier 2 end-to-end exploit
Learn first. capability-evaluation
K3 adaptation. K3 reports a two-tier suite and explicit lower-bound qualification.
Common misconception. Finding a bug is equivalent to completing an exploit chain.
Boundary. Capability is sensitive to target coverage safeguards and evaluator access.
Sources. Kimi K3: Open Frontier Intelligence
Evaluation baselines
Intuition. A score becomes informative only beside a relevant comparator.
Definition. Baselines are models or systems evaluated as reference points under a stated protocol.
Formula or process. gap = score(candidate) - score(baseline)
Learn first. benchmarking
K3 adaptation. K3 is compared with proprietary and open-weight frontier models.
Common misconception. A baseline name alone does not guarantee a fair comparison.
Boundary. Harnesses and effort settings can differ across baselines.
Sources. Kimi K3: Open Frontier Intelligence
Evaluation limitations
Intuition. State what a measurement cannot support.
Definition. Limitations identify uncertainty confounds missing coverage and extrapolations beyond the observed protocol.
Formula or process. supported claim = observation - confounds - extrapolation
Learn first. benchmark-interpretation
K3 adaptation. Course 5A labels paper results local analysis and causal hypotheses separately.
Common misconception. A caveat makes the measurement useless.
Boundary. Unknown unknowns remain outside any declared boundary.
Sources. Kimi K3: Open Frontier Intelligence
Evaluation protocols
Intuition. The recipe is part of the result.
Definition. A protocol fixes tasks metrics prompts sampling tools harnesses budgets and aggregation rules.
Formula or process. result = f(model, tasks, harness, sampling, tools)
Learn first. benchmarking
K3 adaptation. Section 6 records task-specific top-p values tools harnesses and repeated runs.
Common misconception. A benchmark score is not a model-only property.
Boundary. Published protocols may omit operational details needed for exact reproduction.
Sources. Kimi K3: Open Frontier Intelligence
Inference cost
Intuition. Capability must be considered beside resources consumed.
Definition. Per-task inference cost aggregates priced tokens compute or service charges required by the evaluation protocol.
Formula or process. cost/task = sum(input + output + tool costs)
Learn first. evaluation-protocols
K3 adaptation. K3 compares score against per-task cost on four coding and agentic suites.
Common misconception. Token price alone equals total deployment cost.
Boundary. Provider prices harness overhead and cache policy change the estimate.
Sources. Kimi K3: Open Frontier Intelligence
Internal evaluation
Intuition. Track product-specific failures that public tests miss.
Definition. An internal evaluation is maintained by the model developer and can be refreshed around emerging use cases and failure modes.
Formula or process. failures_t -> eval_{t+1} -> training_{t+1}
Learn first. benchmark-design
K3 adaptation. Kimi uses coding general-agent and conversational suites to guide iteration.
Common misconception. Internal automatically means invalid or unbiased.
Boundary. Private tasks and rubrics reduce external auditability.
Sources. Kimi K3: Open Frontier Intelligence
Leaderboards
Intuition. A rank is a moving snapshot of one metric.
Definition. A leaderboard orders submissions under a shared scoring process at a given time.
Formula or process. rank_t = order(scores_t)
Learn first. comparative-evaluation
K3 adaptation. Section 6 dates all cited rankings because entries and Elo values drift.
Common misconception. Rank is an intrinsic permanent property of a model.
Boundary. New entries votes and policy changes alter rank.
Sources. Kimi K3: Open Frontier Intelligence
Reproducibility
Intuition. Another evaluator should be able to reconstruct the claim.
Definition. Reproducibility requires stable data versions configurations code seeds run counts and raw outputs.
Formula or process. claim -> data + config + code + outputs
Learn first. evaluation-protocols
K3 adaptation. Section 6 records dates task branches hardware and leaderboard sources for many results.
Common misconception. Naming the benchmark is sufficient provenance.
Boundary. Third-party services and drifting leaderboards may not be replayable.
Sources. Kimi K3: Open Frontier Intelligence
Result provenance
Intuition. Know who ran the evaluation and where every number came from.
Definition. Provenance records the source date method and transformation chain behind a result.
Formula or process. result <- source <- protocol <- run
Learn first. reproducibility
K3 adaptation. K3 separates internal runs cited leaderboards and independent evaluations.
Common misconception. A number copied into a table becomes locally verified.
Boundary. Provenance does not by itself prove methodological quality.
Sources. Kimi K3: Open Frontier Intelligence
Sampling parameters
Intuition. Decoding choices alter the behavior being measured.
Definition. Temperature top-p and sample count control randomness diversity and aggregation during generation.
Formula or process. p'(x) proportional to exp(log p(x)/T) within top-p
Learn first. None
K3 adaptation. K3 uses temperature 1.0 with top-p selected by task type.
Common misconception. Sampling is an implementation detail unrelated to score.
Boundary. Stochastic results require repeated runs and uncertainty reporting.
Sources. Kimi K3: Open Frontier Intelligence
Third-party evaluation
Intuition. Independent measurement reduces reliance on one evaluator.
Definition. An external organization selects or runs the protocol and publishes the resulting comparison.
Formula or process. confidence increases with independent methods and sources
Learn first. provenance, comparative-evaluation
K3 adaptation. Artificial Analysis Vals AI and arena rankings triangulate K3 from different perspectives.
Common misconception. Independence guarantees an unbiased or stable metric.
Boundary. Leaderboards retain their own judges pricing and sampling choices.
Sources. Kimi K3: Open Frontier Intelligence
Foundations
Activation functions
Intuition. Introduce nonlinear transformations between learned linear maps.
Definition. An element-wise nonlinear function changes representation geometry and model expressivity.
Formula or process. y=f(Wx)
Learn first. None
K3 adaptation. Stable LatentMoE chooses bounded SiTU branches for numerical control.
Common misconception. A nonlinear function alone does not create gating.
Boundary. Stability depends on the surrounding network and precision.
Sources. Attention Is All You Need
Language-model scaling
Intuition. Increase model data and compute while tracking predictable changes in loss and capability.
Definition. Empirical scaling studies relate training resources and model size to held-out loss under a specified regime.
Formula or process. L(N,D)=L_inf+A N^-alpha+B D^-beta
Learn first. transformer-basics
K3 adaptation. K3 combines a 2.8T-parameter sparse model with a reported 3T-class training system.
Common misconception. Parameter count alone does not specify training compute or inference cost.
Boundary. Extrapolation depends on data quality architecture and fitted range.
Sources. Scaling Laws for Neural Language Models
Model scaling
Intuition. Grow capacity while controlling active computation and systems cost.
Definition. Choose depth width expert count and sparsity jointly with data and compute budgets.
Formula or process. capacity != active FLOPs != training compute
Learn first. llm-scaling
K3 adaptation. K3 reports 2.8T total parameters with 104B active parameters.
Common misconception. Total parameters are not activated on every token.
Boundary. Reported scale does not by itself imply capability.
Sources. Kimi K3: Open Frontier Intelligence
Transformer basics
Intuition. Alternate information mixing with token-wise transformation under residual paths.
Definition. A causal Transformer stacks attention feed-forward normalization and residual operations to predict the next token.
Formula or process. h -> attention -> residual -> FFN -> residual
Learn first. attention-mechanism
K3 adaptation. K3 replaces or modifies each major routing axis while retaining causal next-token prediction.
Common misconception. Transformer does not mean softmax attention at every layer.
Boundary. This prerequisite omits training and systems details.
Sources. Attention Is All You Need
Infrastructure
Activation memory
Intuition. Intermediate tensors occupy memory until backward no longer needs them.
Definition. Forward values have lifetimes determined by backward dependencies checkpointing and recomputation.
Formula or process. peak = max_t sum_i size_i live_i(t)
Learn first. distributed-training
K3 adaptation. K3 manages activation lifetime and remote offload across pipeline ranks.
Common misconception. Parameter memory is not the entire training footprint.
Boundary. Policies trade memory compute and communication.
Sources. Kimi K3: Open Frontier Intelligence
Activation recomputation
Intuition. Recalculate cheap intermediates instead of storing them.
Definition. Checkpoint selected values and replay forward regions during backward to reduce memory.
Formula or process. less memory <-> more FLOPs
Learn first. activation-memory
K3 adaptation. Element-wise and selected training intermediates are recomputed.
Common misconception. Full-layer checkpointing is not the only granularity.
Boundary. Excess replay can extend the critical path.
Sources. Kimi K3: Open Frontier Intelligence
Adaptive throttling
Intuition. Reduce or delay work when a shared resource approaches unsafe load.
Definition. Feedback from queue cache or memory pressure changes admission or concurrency.
Formula or process. concurrency_t+1 = policy(load_t)
Learn first. distributed-systems
K3 adaptation. Long-context rollout pressure controls active generation.
Common misconception. Fixed concurrency is not adaptive control.
Boundary. Delayed feedback can oscillate.
Sources. Kimi K3: Open Frontier Intelligence
All-to-all communication
Intuition. Every rank may send different data to every other rank.
Definition. A collective redistributes partitioned buffers among all participants.
Formula or process. rank_i sends block_ij to rank_j
Learn first. distributed-systems
K3 adaptation. Expert dispatch and combine require structured rank exchange.
Common misconception. Equal total bytes do not imply equal completion time.
Boundary. Topology and message size dominate efficiency.
Sources. Kimi K3: Open Frontier Intelligence
Communication overlap
Intuition. Transfer data while independent computation proceeds.
Definition. Dependency-aware streams hide communication behind compute when resources and timing permit.
Formula or process. visible time = max(compute, communication) when fully overlapped
Learn first. distributed-systems
K3 adaptation. K3 pipelines expert and optimizer communication with kernels.
Common misconception. Concurrent APIs do not guarantee physical overlap.
Boundary. Short compute windows expose transfer time.
Sources. Kimi K3: Open Frontier Intelligence
Consistent hashing
Intuition. Map keys to a changing fleet while minimizing remapping.
Definition. A hash ring or related scheme assigns keys to nodes with limited movement when membership changes.
Formula or process. node = successor(hash(key))
Learn first. distributed-systems
K3 adaptation. Prefix affinity can guide requests toward resident state.
Common misconception. Consistent hashing alone does not manage load.
Boundary. Hot keys still create hotspots.
Sources. Kimi K3: Open Frontier Intelligence
Context parallelism
Intuition. Split one sequence across workers.
Definition. Devices own different sequence segments and exchange summaries or attention data required for exact results.
Formula or process. sequence = concat(segment_1 ... segment_P)
Learn first. distributed-training
K3 adaptation. KCP composes affine KDA segment transforms.
Common misconception. Ring metadata is not distributed context execution.
Boundary. Communication pattern depends on the attention operator.
Sources. Kimi K3: Open Frontier Intelligence, Ring Attention with Blockwise Transformers for Near-Infinite Context
Copy on write
Intuition. Share unchanged state and copy a page only when a branch modifies it.
Definition. Multiple logical versions reference common storage until a write creates a private copy.
Formula or process. fork cost proportional to changed pages
Learn first. None
K3 adaptation. AgentENV forks worlds without eagerly cloning them.
Common misconception. Forking does not mean zero future storage.
Boundary. Write-heavy branches lose sharing.
Sources. Kimi K3: Open Frontier Intelligence
Critical path analysis
Intuition. Only the longest dependency chain determines completion time.
Definition. In a directed task graph the critical path is the maximum-duration path to completion.
Formula or process. latency = max_path sum duration
Learn first. None
K3 adaptation. Vision encoder work is moved into non-critical pipeline intervals.
Common misconception. Faster off-path work may not reduce step time.
Boundary. Durations change with contention.
Sources. Kimi K3: Open Frontier Intelligence
cuBLAS
Intuition. Use NVIDIA's tuned dense linear-algebra kernels as a production baseline.
Definition. cuBLAS supplies GPU BLAS operations including highly optimized matrix multiplication for supported shapes and precisions.
Formula or process. C = alpha A B + beta C
Learn first. gpu-kernels, tensor-cores
K3 adaptation. Custom MiniTriton kernels are justified only where specialized operators or fusion exceed generic dense-GEMM needs.
Common misconception. A custom kernel is not automatically faster than cuBLAS.
Boundary. Performance is shape hardware precision and layout specific.
Sources. NVIDIA cuBLAS Library
Data parallelism
Intuition. Give replicas different samples and combine their gradients.
Definition. Each replica owns model parameters and processes a different microbatch before synchronizing the update.
Formula or process. g = sum_r g_r / R
Learn first. distributed-training
K3 adaptation. K3 combines data parallelism with parameter and gradient sharding.
Common misconception. Replication does not reduce per-replica model-state memory.
Boundary. Gradient communication grows with model size.
Sources. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Decoupled Encoder Process
Intuition. Let vision encoding advance independently so variable image work stops gating the language pipeline.
Definition. Encoder tasks are scheduled and buffered separately from language-model stages then joined at a defined feature boundary.
Formula or process. image queue -> encoder workers -> feature boundary -> LM
Learn first. multimodal-encoders, pipeline-parallelism, critical-path-analysis
K3 adaptation. K3 combines dynamic context groups and pipeline bubble filling around the vision encoder.
Common misconception. Decoupling does not change the model's required feature dependency.
Boundary. Queueing and staleness bounds are paper-system details.
Sources. Kimi K3: Open Frontier Intelligence
Distributed systems
Intuition. Coordinate computation and state across machines that communicate and fail independently.
Definition. A distributed system preserves required semantics while work and data span multiple processes or nodes.
Formula or process. global result = compose(local work, communication)
Learn first. None
K3 adaptation. K3 coordinates training rollout and serving state across accelerator fleets.
Common misconception. More machines do not automatically create useful parallelism.
Boundary. Communication stragglers and failures shape behavior.
Sources. Kimi K3: Open Frontier Intelligence
Distributed training
Intuition. Split model data and optimizer work across devices without changing the intended update.
Definition. Parallel workers coordinate forward backward communication and parameter updates.
Formula or process. step = forward + backward + synchronize + update
Learn first. distributed-systems
K3 adaptation. Combines expert pipeline tensor and context parallelism.
Common misconception. Device count is not equivalent to scaling efficiency.
Boundary. Numerical order and communication affect results.
Sources. Kimi K3: Open Frontier Intelligence
Expert parallelism
Intuition. Place different experts on different devices and route tokens between them.
Definition. Sparse MoE execution dispatches token representations to devices owning selected experts and combines results.
Formula or process. token -> selected expert ranks -> combine
Learn first. distributed-training, moe-routing
K3 adaptation. MoonEP balances routed work through bounded redundancy.
Common misconception. Sparse activation does not guarantee balanced ranks.
Boundary. All-to-all traffic can dominate.
Sources. Kimi K3: Open Frontier Intelligence
External model-state pool
Intuition. Preserve a long trajectory without reserving its rollout GPU.
Definition. KV blocks and recurrent checkpoints move to a separately managed pool and return with identity and causal-boundary metadata.
Formula or process. trajectory id -> {MLA KV, KDA state, boundary}
Learn first. rollout-systems, prefix-caching
K3 adaptation. K3 externalizes million-token rollout state during pauses.
Common misconception. Offloading state is not equivalent to discarding and recomputing it.
Boundary. No distributed pool or bandwidth result is reproduced locally.
Sources. Kimi K3: Open Frontier Intelligence
GPU kernels
Intuition. Implement numerical operators directly for accelerator hardware.
Definition. Parallel programs map tensor computation to threads memory and device instructions.
Formula or process. operator -> GPU program
Learn first. None
K3 adaptation. Forms a verified RL task family.
Common misconception. Correctness and speed are separate gates.
Boundary. Performance is hardware-specific.
Sources. Kimi K3: Open Frontier Intelligence
GPU scheduling
Intuition. Order kernels and transfers so scarce device resources remain busy.
Definition. Streams dependencies priorities and launch shapes determine when GPU operations execute.
Formula or process. makespan = critical path of kernels and transfers
Learn first. gpu-kernels
K3 adaptation. MoonEP overlaps dispatch prefetch and expert GEMMs.
Common misconception. Asynchronous launch does not guarantee overlap.
Boundary. Resource contention can serialize operations.
Sources. Kimi K3: Open Frontier Intelligence
Gradient-buffer reuse
Intuition. Reuse already reserved gradient storage as temporary forwarding workspace when gradients are not live.
Definition. A lifetime-aware runtime aliases non-overlapping tensor roles to the same physical buffer.
Formula or process. if live(A) and live(B) do not overlap then storage(A) may equal storage(B)
Learn first. unified-activation-manager, activation-memory
K3 adaptation. K3 reuses gradient buffers for non-policy model forwarding.
Common misconception. Safe aliasing requires proven non-overlapping lifetimes.
Boundary. Local code does not implement allocator-level aliasing.
Sources. Kimi K3: Open Frontier Intelligence
Interleaved 1F1B
Intuition. Alternate one forward and one backward microbatch after pipeline warm-up.
Definition. A pipeline schedule keeps stages busy by interleaving forward and backward work across virtual stages while respecting dependencies.
Formula or process. warmup -> alternate 1F and 1B -> drain
Learn first. virtual-pipeline-stages
K3 adaptation. K3 balances unequal activation lifetimes across ranks under interleaved 1F1B.
Common misconception. Every rank does not hold the same number of live activations.
Boundary. Warm-up and placement create rank-dependent memory peaks.
Sources. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Memory pooling
Intuition. Reuse reserved storage instead of repeatedly allocating it.
Definition. A pool manages blocks with known sizes and lifetimes for multiple operations.
Formula or process. reserve once -> assign -> release -> reuse
Learn first. activation-memory
K3 adaptation. Unified activation management supports plug-in policies and buffer reuse.
Common misconception. Preallocation is not automatically zero-copy.
Boundary. Fragmentation can remain inside the pool.
Sources. Kimi K3: Open Frontier Intelligence
MicroVMs
Intuition. Use lightweight virtual machines for strong isolation with fast lifecycle operations.
Definition. A minimal virtual machine packages a guest kernel and constrained devices for one workload.
Formula or process. isolated guest + snapshot lifecycle
Learn first. sandboxing
K3 adaptation. AgentENV reports resumable high-density environments.
Common misconception. A process or dictionary is not a microVM.
Boundary. Startup snapshot and storage costs remain.
Sources. Kimi K3: Open Frontier Intelligence
MiniTriton
Intuition. Express specialized K3 kernels above raw CUDA while retaining control over layouts and fusion.
Definition. K3's reported compiler and DSL lower operator programs through compiler IR to GPU code and integrate runtime autograd and communication concerns.
Formula or process. DSL -> compiler IR -> GPU program
Learn first. gpu-kernels, cublas
K3 adaptation. K3 uses MiniTriton for its long tail of architecture-specific kernels.
Common misconception. MiniTriton is not evidence that every operator should replace vendor libraries.
Boundary. The compiler is discussed as paper-reported infrastructure and is not implemented here.
Sources. Kimi K3: Open Frontier Intelligence
OverlayBD
Intuition. Share a common read-only environment image while each agent records only its private changes.
Definition. A layered block-device design composes immutable base data with copy-on-write overlays for fast fork snapshot and recovery.
Formula or process. world = base image + trajectory overlay
Learn first. copy-on-write, microvms
K3 adaptation. AgentENV uses layered storage to make persistent agent worlds denser.
Common misconception. Overlay storage provides sharing but is not by itself sandbox isolation.
Boundary. The repository contains no OverlayBD integration.
Sources. Kimi K3: Open Frontier Intelligence, DADI: Block-Level Image Service for Agile and Elastic Application Deployment
P2P-based Muon orthogonalization
Intuition. Borrow peer memory and links for an optimizer transform too large to stage comfortably on one rank.
Definition. Parameter or momentum chunks move peer-to-peer to workers that execute Newton-Schulz orthogonalization and return transformed updates while other work proceeds.
Formula or process. chunk -> peer NS transform -> return update
Learn first. muon, communication-overlap
K3 adaptation. K3 reports P2P orchestration for Muon in its memory-efficient training stack.
Common misconception. P2P changes where Muon runs not the Newton-Schulz mathematics.
Boundary. No local distributed optimizer reproduces the schedule.
Sources. Kimi K3: Open Frontier Intelligence, Muon: An optimizer for hidden layers in neural networks
Performance profiling
Intuition. Measure where execution time and resources go.
Definition. Instrument runtime to attribute latency throughput and utilization to program regions.
Formula or process. time=sum component time
Learn first. None
K3 adaptation. Supports kernel optimization reward.
Common misconception. One timing sample is not a benchmark.
Boundary. Measurement perturbs execution.
Sources. Kimi K3: Open Frontier Intelligence
Pipeline parallelism
Intuition. Place consecutive model stages on different devices and stream microbatches through them.
Definition. A schedule overlaps forward and backward stage work while respecting dependencies.
Formula or process. microbatch -> stage_1 -> ... -> stage_P
Learn first. distributed-training
K3 adaptation. ViT work and activation offload are coordinated with LM pipeline bubbles.
Common misconception. Pipeline bubbles are not always wasted if other work can fill them.
Boundary. Stage imbalance sets throughput.
Sources. Kimi K3: Open Frontier Intelligence
Pipeline ZeRO-2
Intuition. Move gradient shards through the pipeline instead of keeping a complete gradient copy on every rank.
Definition. Gradient ownership and transfer are coordinated with pipeline execution so shards can be reduced offloaded and consumed near their update window.
Formula or process. gradient bytes per owner approximately gradient bytes / shard group
Learn first. zero-one, pipeline-parallelism, one-f-one-b
K3 adaptation. K3 reports Pipeline ZeRO-2 gradient sharding and offloading.
Common misconception. It is not simply ordinary ZeRO-2 running independently of the pipeline schedule.
Boundary. The paper-level description does not expose a reproducible local implementation.
Sources. Kimi K3: Open Frontier Intelligence, ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Rollout systems
Intuition. Generate policy trajectories while managing workers queues environments and state.
Definition. A rollout runtime schedules inference action execution observations persistence and handoff to optimization.
Formula or process. policy -> action -> environment -> observation -> queue
Learn first. reinforcement-learning
K3 adaptation. Partial rollouts and external state pools preserve long jobs.
Common misconception. Rollout generation is not a stateless batch job.
Boundary. Policy staleness and tail latency matter.
Sources. Kimi K3: Open Frontier Intelligence
Sandboxing
Intuition. Let untrusted actions run inside isolated resource boundaries.
Definition. A sandbox controls filesystem process network and lifecycle access for a workload.
Formula or process. action inside policy-constrained environment
Learn first. agentic-environments
K3 adaptation. AgentENV supplies isolated persistent worlds.
Common misconception. A Python dictionary is not isolation.
Boundary. Fidelity and security compete with density.
Sources. Kimi K3: Open Frontier Intelligence
Tensor cores
Intuition. Execute dense matrix fragments at high throughput.
Definition. Accelerator units perform specialized mixed-precision matrix multiply-accumulate operations.
Formula or process. D = A B + C
Learn first. gpu-kernels
K3 adaptation. FlashKDA schedules chunk work to expose matrix operations.
Common misconception. Tensor-core peak does not imply application utilization.
Boundary. Layout precision and tile shape constrain use.
Sources. Kimi K3: Open Frontier Intelligence
Training checkpointing
Intuition. Persist enough distributed state to resume a long run safely.
Definition. Periodic snapshots capture parameters, optimizer state, scheduler state, and progress metadata.
Formula or process. checkpoint=(weights, optimizer, schedule, data position)
Learn first. training-recipe
K3 adaptation. Checkpoint policy belongs to the coordinated pretraining program.
Common misconception. Saving model weights alone is not a fully resumable checkpoint.
Boundary. Recovery guarantees depend on distributed storage and consistency design.
Sources. Kimi K3: Open Frontier Intelligence
Unified activation manager
Intuition. Decide tensor placement from lifetime rather than from operator ownership.
Definition. A common runtime tracks when activations are created consumed recomputed offloaded released and reused across model components.
Formula or process. peak = max_t sum_i size_i live_i(t)
Learn first. activation-memory, memory-pooling, recomputation
K3 adaptation. K3 combines component-specific policies under one activation lifecycle.
Common misconception. One universal checkpoint interval is not a unified manager.
Boundary. The production policy and allocator are not reproduced locally.
Sources. Kimi K3: Open Frontier Intelligence
Virtual pipeline stages
Intuition. Give each physical pipeline rank several non-adjacent model chunks so work can interleave.
Definition. A physical device cycles through multiple logical stages to reduce idle bubbles and improve balance.
Formula or process. physical rank owns virtual stages v_1 ... v_k
Learn first. pipeline-parallelism
K3 adaptation. K3 reports pipeline parallelism with virtual stages for 3T-class training.
Common misconception. A virtual stage is scheduling and placement not another GPU.
Boundary. More interleaving increases activation and communication complexity.
Sources. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
ZeRO-1 data parallelism
Intuition. Replicate weights and gradients but partition optimizer state.
Definition. Data-parallel ranks own disjoint optimizer-state shards and update their parameter partitions before synchronizing weights.
Formula or process. optimizer bytes per rank approximately optimizer bytes / R
Learn first. data-parallelism
K3 adaptation. ZeRO-1 is one layer in K3's multidimensional pretraining layout.
Common misconception. ZeRO-1 does not shard parameters or gradients during ordinary storage.
Boundary. Weight and gradient memory remain replicated.
Sources. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Zero-copy communication
Intuition. Put arriving tokens directly where the consumer kernel expects them.
Definition. Preplanned offsets and registered buffers avoid an intermediate receive-then-repack copy before expert computation.
Formula or process. network -> expert-grouped destination -> GEMM
Learn first. expert-parallelism, all-to-all-communication
K3 adaptation. Perfect MoonEP balance makes destination shapes and offsets static.
Common misconception. Preallocating a tensor and calling copy is not network zero-copy.
Boundary. Hardware transport registration and topology determine feasibility.
Sources. Kimi K3: Open Frontier Intelligence
Latent Attention
Gated MLA with NoPE
Intuition. Periodically open a compressed global content lookup and gate its output.
Definition. Global MLA output is modulated by a learned full-rank sigmoid gate before projection.
Formula or process. y_t=W_o[sigma(W_g x_t) odot MLA(x_t)]
Learn first. deepseek-mla, nope
K3 adaptation. One Gated MLA follows every three KDA layers, plus one final Gated MLA.
Common misconception. NoPE does not mean the full hybrid network has no order signal.
Boundary. Repository evidence validates a miniature interface, not paper-scale cache economics.
Sources. Kimi K3: Open Frontier Intelligence
Low-rank latent compression
Intuition. Preserve a compact coordinate system and reconstruct wider representations only when used.
Definition. A learned down-projection maps a hidden vector into a smaller latent that can be cached or routed efficiently.
Formula or process. c_t = W_c x_t, dim(c_t) << dim(x_t)
Learn first. None
K3 adaptation. Used by Gated MLA and routed LatentMoE experts.
Common misconception. Low rank reduces representation width; it does not make global attention recurrent.
Boundary. Compression quality is learned and task-dependent.
Sources. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Multi-Head Latent Attention
Intuition. Cache one compact latent per token instead of full keys and values.
Definition. MLA compresses hidden states into KV latents and reconstructs head-specific attention representations.
Formula or process. c_t=W_c x_t; K_t,V_t=up_project(c_t)
Learn first. softmax-attention, kv-cache, low-rank-compression
K3 adaptation. K3 adds a full-rank output gate and removes explicit positional encoding in its Gated MLA layers.
Common misconception. MLA still maintains token-indexed latent cache entries.
Boundary. Exact cache savings depend on model dimensions and implementation.
Sources. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
No positional encoding
Intuition. Do not inject an explicit RoPE or absolute-position transform into the MLA query/key path.
Definition. Queries and keys are projected without an explicit positional embedding operation.
Formula or process. q_t=W_q x_t; k_t=W_k c_t
Learn first. softmax-attention
K3 adaptation. Interleaved recurrent KDA layers carry causal order and recency information between MLA layers.
Common misconception. NoPE is not the same as position-insensitive computation.
Boundary. The repository does not isolate every source of learned positional behavior.
Sources. Kimi K3: Open Frontier Intelligence
Mathematics
Affine transforms
Intuition. Represent an update as a linear map plus an offset.
Definition. An affine transform maps x to Mx plus b and composes associatively as a pair.
Formula or process. (M2,b2) circle (M1,b1) = (M2M1,M2b1+b2)
Learn first. None
K3 adaptation. Each KDA segment is summarized as an incoming-state transform.
Common misconception. Affine is not necessarily scalar or diagonal.
Boundary. Exact form depends on the recurrence.
Sources. Kimi K3: Open Frontier Intelligence
Quantiles
Intuition. Locate a threshold by rank rather than by an assumed distribution.
Definition. The q-quantile is a value at or below which approximately fraction q of observations lie.
Formula or process. Q(q)=inf{x:F(x)>=q}
Learn first. None
K3 adaptation. Routing score-margin quantiles determine expert-specific selection biases.
Common misconception. A quantile is not generally the mean.
Boundary. Estimates are noisy with small or shifting samples.
Sources. Kimi K3: Open Frontier Intelligence
Multimodality
Multimodal pretraining data
Intuition. Pair or interleave perceptual evidence with language and structured supervision.
Definition. Images, video, text, coordinates, and generated render-code pairs are serialized as training examples.
Formula or process. sequence = interleave(text tokens, visual tokens, coordinates)
Learn first. vision-transformer
K3 adaptation. K3 includes captions, OCR, perception, video, visual coding, and programmatic pairs.
Common misconception. Multimodal data is broader than image-caption pairs.
Boundary. Data composition alone does not establish grounded reasoning.
Sources. Kimi K3: Open Frontier Intelligence
Multimodal tokenization
Intuition. Express different modalities as one ordered sequence of model-width vectors.
Definition. Modality-specific encoders and projectors convert text images or video into embeddings placed in causal order.
Formula or process. modal input -> encoder -> projector -> shared token stream
Learn first. transformer-basics
K3 adaptation. MoonViT-V2 features join text before the shared K3 backbone.
Common misconception. Pixels do not become text IDs directly.
Boundary. Encoder and projector quality bound the shared representation.
Sources. Kimi K3: Open Frontier Intelligence
Vision Transformer patchification
Intuition. Turn spatial image regions into a token sequence.
Definition. Split an image into patches, project each patch, and process them with Transformer layers.
Formula or process. patches(image) -> linear projection -> visual token sequence
Learn first. softmax-attention
K3 adaptation. MoonViT-V2 is trained natively with the language model and projected into its embedding width.
Common misconception. The repository patchifier is not a full MoonViT-V2 reproduction.
Boundary. Miniature output validates shape and flow, not vision benchmark quality.
Sources. An Image is Worth 16x16 Words
Visual tokenization
Intuition. Convert spatial regions into a sequence while retaining layout information.
Definition. Images or frames are partitioned or pooled into visual units encoded and projected as vectors.
Formula or process. H x W image -> N visual vectors
Learn first. vision-transformer
K3 adaptation. MoonViT-V2 supplies native visual features to K3.
Common misconception. Patchification alone is not a trained vision encoder.
Boundary. Resolution and merging policy determine token count and detail.
Sources. An Image is Worth 16x16 Words
Normalization
RMSNorm
Intuition. Stabilize vector scale without subtracting its mean.
Definition. Divide a vector by its root-mean-square magnitude and apply a learned scale.
Formula or process. RMSNorm(x)=g odot x/sqrt(mean(x^2)+epsilon)
Learn first. None
K3 adaptation. Used in depth routing and stable latent expert paths.
Common misconception. Normalization controls scale but does not itself balance experts.
Boundary. Placement and precision affect behavior.
Sources. Root Mean Square Layer Normalization
Optimization
Curriculum learning
Intuition. Change task difficulty or constraints as learning proceeds.
Definition. Training presents a staged distribution rather than a stationary sample.
Formula or process. P_t(task)
Learn first. None
K3 adaptation. Anneals effort-budget multipliers.
Common misconception. Curriculum order is a design choice.
Boundary. Poor staging can cause forgetting.
Sources. Kimi K3: Open Frontier Intelligence
Gradient descent
Intuition. Change parameters in a direction that locally reduces loss.
Definition. Compute derivatives of loss with respect to parameters and apply a scaled update possibly with optimizer state.
Formula or process. theta_{t+1}=theta_t-eta grad L(theta_t)
Learn first. None
K3 adaptation. Per-Head Muon transforms matrix-valued momentum before applying the update.
Common misconception. The raw negative gradient is not every optimizer's final update.
Boundary. Local derivatives do not guarantee global convergence.
Sources. Kimi K3: Open Frontier Intelligence
Matrix orthogonalization
Intuition. Equalize update strength across independent singular directions.
Definition. Map a matrix toward its polar factor or a scaled semi-orthogonal matrix subject to its rectangular shape.
Formula or process. G=U Sigma V^T -> UV^T
Learn first. matrix-optimization
K3 adaptation. Newton-Schulz approximates the transform on each head-local block.
Common misconception. Wide matrices cannot have orthonormal columns.
Boundary. Rank deficiency scaling and finite iterations affect the result.
Sources. Muon: An optimizer for hidden layers in neural networks
Matrix-valued optimization
Intuition. Treat a parameter update as geometry over rows and columns rather than unrelated scalars.
Definition. An optimizer transforms matrix gradients or momentum using operations that couple their singular directions.
Formula or process. Delta W=T(M)
Learn first. gradient-descent
K3 adaptation. Attention projection updates are partitioned by head before transformation.
Common misconception. Reshaping arbitrary tensors into matrices is not geometry-neutral.
Boundary. The chosen partition changes the update.
Sources. Muon: An optimizer for hidden layers in neural networks
Mixed-precision training
Intuition. Use compact number formats where safe while protecting sensitive arithmetic.
Definition. Training assigns different numeric precisions to tensors and operations to balance throughput and stability.
Formula or process. compute_low_precision + selected_high_precision_accumulation
Learn first. None
K3 adaptation. K3's recipe couples precision controls with clipping, QB, and Per-Head Muon.
Common misconception. Lower precision is not merely storage compression.
Boundary. Stability is hardware-, kernel-, and scale-dependent.
Sources. Kimi K3: Open Frontier Intelligence
Muon optimizer
Intuition. Orthogonalize matrix momentum before using it as a parameter update.
Definition. Muon forms momentum and applies an approximate polar-factor transform to hidden-layer matrix updates.
Formula or process. M_t=mu M_{t-1}+g_t; Delta W=NS(M_t)
Learn first. gradient-descent, matrix-optimization, orthogonalization
K3 adaptation. K3 applies Muon separately to attention heads.
Common misconception. Muon does not replace the optimizer for every scalar vector or embedding parameter.
Boundary. Benefits depend on architecture scale hyperparameters and implementation.
Sources. Muon: An optimizer for hidden layers in neural networks
MXFP4
Intuition. Store expert weights using a compact microscaling format.
Definition. A four-bit floating representation shares scale metadata across value blocks.
Formula or process. 4-bit values plus block scale
Learn first. mixed-precision-training
K3 adaptation. Quantizes routed expert weights.
Common misconception. Not all K3 weights use MXFP4.
Boundary. Error depends on scaling and distribution.
Sources. Kimi K3: Open Frontier Intelligence, Microscaling Data Formats for Deep Learning
MXFP8 activation format
Intuition. Use block-scaled eight-bit floating values for expert inputs while retaining more sensitive paths at higher precision.
Definition. An eight-bit microscaling floating format represents values alongside a scale shared by a block.
Formula or process. 8-bit values + shared block scale
Learn first. mixed-precision-training
K3 adaptation. K3 pairs MXFP8 routed-expert input activations with MXFP4 expert weights during SFT RL rollout and training.
Common misconception. K3 does not quantize every activation to MXFP8.
Boundary. Accuracy and speed depend on exact kernels scaling blocks and hardware support.
Sources. Kimi K3: Open Frontier Intelligence, Microscaling Data Formats for Deep Learning
Newton–Schulz orthogonalization
Intuition. Repeated matrix multiplications push singular values toward a shared orthogonal scale.
Definition. A polynomial matrix iteration approximates an orthogonalized update without explicit SVD.
Formula or process. X_{k+1}=aX_k+bX_kX_k^TX_k+cX_k(X_k^TX_k)^2
Learn first. None
K3 adaptation. Applied independently to per-head projection updates.
Common misconception. It orthogonalizes update directions, not model activations.
Boundary. Convergence depends on scaling, coefficients, precision, and iteration count.
Sources. Muon: An optimizer for hidden layers in neural networks
Optimization schedules
Intuition. Change step size over time to enter, traverse, and settle the loss landscape.
Definition. Warmup and decay rules map training progress to learning rate.
Formula or process. eta(t)=warmup(t) then cosine_decay(t)
Learn first. None
K3 adaptation. K3 reports cosine decay with 1% warmup and independent schedule tuning.
Common misconception. Schedules cannot be compared fairly without retuning other controls.
Boundary. Schedule arithmetic does not guarantee stable optimization.
Sources. Kimi K3: Open Frontier Intelligence
Optimizer state
Intuition. Training retains update statistics in addition to weights and gradients.
Definition. Momentum variance or matrix statistics persist across steps and consume storage and bandwidth.
Formula or process. state_t+1 = update(state_t, gradient_t)
Learn first. training-recipe
K3 adaptation. Per-Head Muon buffers are gathered and transformed in chunks.
Common misconception. Low-precision weights do not remove optimizer memory.
Boundary. Sharding adds communication.
Sources. Kimi K3: Open Frontier Intelligence
Per-Head Muon
Intuition. Preserve head-local update geometry instead of orthogonalizing one concatenated projection matrix.
Definition. Matrix-valued optimizer updates are partitioned by attention head and orthogonalized independently.
Formula or process. Delta W_h = NS(optimizer_momentum_h) for each head h
Learn first. newton-schulz
K3 adaptation. K3 applies Muon at attention-head granularity.
Common misconception. Muon is not a replacement for every optimizer state or parameter group.
Boundary. Repository tests validate mechanics, not paper-scale convergence gains.
Sources. Kimi K3: Open Frontier Intelligence, Muon: An optimizer for hidden layers in neural networks
Quantization-aware training
Intuition. Train while simulating lower-precision arithmetic.
Definition. Forward computation exposes quantization error so updates can adapt parameters to it.
Formula or process. forward Q(w); update w
Learn first. None
K3 adaptation. Runs throughout SFT and RL.
Common misconception. It is not post-hoc rounding.
Boundary. Simulation must match deployed kernels.
Sources. Kimi K3: Open Frontier Intelligence, Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Training stability
Intuition. Keep interacting signals inside a regime where useful learning continues.
Definition. Clipping, bounded activations, balancing, precision policy, and schedule controls limit destructive excursions.
Formula or process. stable step = controlled geometry + magnitude + routing + precision
Learn first. training-recipe
K3 adaptation. K3 combines Per-Head Muon, weight clipping, SiTU-GLU, and Quantile Balancing.
Common misconception. No single scalar metric owns stability.
Boundary. Small tests cannot certify trillion-parameter behavior.
Sources. Kimi K3: Open Frontier Intelligence
Posttraining
Agentic environments
Intuition. Let a policy change and observe an external world.
Definition. Stateful tool interfaces produce observations and verifiable outcomes over multiple steps.
Formula or process. s,a -> s',o
Learn first. reinforcement-learning
K3 adaptation. Provides long-horizon task worlds.
Common misconception. A chat transcript alone is not a world state.
Boundary. Fidelity and evaluation remain difficult.
Sources. Kimi K3: Open Frontier Intelligence
Agentic Generative Reward Model
Intuition. Make a judge write the grading logic before it chooses among open-ended outputs.
Definition. A reward-model agent reads the problem and candidates generates a rubric scores each candidate and records the evidence in a scorepad.
Formula or process. problem + candidates -> rubric -> scores -> scorepad
Learn first. reinforcement-learning, verifiable-rewards
K3 adaptation. Supplies feedback for non-verifiable general tasks where deterministic checkers are unavailable.
Common misconception. A generated rubric is not ground truth.
Boundary. The paper does not publish independent calibration of judge accuracy.
Sources. Kimi K3: Open Frontier Intelligence
Cold-start policy
Intuition. Give reinforcement learning a minimally competent policy instead of asking reward alone to discover language and tool use.
Definition. A supervised initialization trained on verified demonstrations before outcome-driven policy improvement.
Formula or process. verified demonstrations -> SFT policy -> RL
Learn first. instruction-data, supervised-fine-tuning
K3 adaptation. K3 builds agentic trajectories with synthesis multi-stage verification human annotation and XTML serialization.
Common misconception. Cold start does not mean random initialization or a tiny dataset.
Boundary. The paper does not disclose dataset size or acceptance rates.
Sources. Kimi K3: Open Frontier Intelligence
Long-context reinforcement learning
Intuition. Improve policies whose trajectories and state span very long contexts.
Definition. Rollout collection optimization and cache management preserve learning signals across extended sequences.
Formula or process. trajectory length up to context limit
Learn first. reinforcement-learning, long-context-training
K3 adaptation. K3 reports million-token agentic RL infrastructure.
Common misconception. A large window alone does not preserve rollout state.
Boundary. Cost and stragglers grow with trajectory length.
Sources. Kimi K3: Open Frontier Intelligence
Multi-Teacher On-Policy Distillation
Intuition. Route each student trajectory to the specialist that matches its domain and effort level.
Definition. The student samples its own tokens while one selected teacher supplies a clipped stop-gradient token-level log-ratio reward.
Formula or process. r_t=clip(stopgrad(log pi_teacher-log pi_student))
Learn first. on-policy-distillation, policy-distillation, reasoning-effort
K3 adaptation. Consolidates nine domain-by-effort policies into one K3 policy.
Common misconception. MOPD does not average all nine teachers for every token.
Boundary. The specialist models and full training dynamics are unavailable locally.
Sources. Kimi K3: Open Frontier Intelligence, On-policy distillation
On-policy distillation
Intuition. Teach on prefixes produced by the student itself.
Definition. A teacher evaluates actions sampled in the current student's contexts.
Formula or process. r_t=log pi_teacher/pi_student
Learn first. reinforcement-learning
K3 adaptation. MOPD routes domain-effort teachers.
Common misconception. It is not offline imitation alone.
Boundary. Requires teacher inference.
Sources. Kimi K3: Open Frontier Intelligence, On-policy distillation
On-policy learning
Intuition. Learn from behavior generated by the current policy.
Definition. Collect trajectories under a current or closely related policy before updating.
Formula or process. y ~ pi_theta
Learn first. reinforcement-learning
K3 adaptation. Student prefixes ground MOPD reward.
Common misconception. Paused rollouts can introduce staleness.
Boundary. Collection is expensive.
Sources. Kimi K3: Open Frontier Intelligence
Partial rollout
Intuition. Start an update after enough trajectories finish without discarding slow unfinished work.
Definition. Once a completion fraction is reached the remaining trajectories are paused with state preserved and resumed in a later iteration.
Formula or process. complete >= lambda batch; pause and resume stragglers
Learn first. on-policy-learning, long-horizon-agents
K3 adaptation. Prevents long-tail agent rollouts from stalling every K3 RL update.
Common misconception. A paused trajectory is not necessarily truncated or assigned failure reward.
Boundary. Resumption introduces policy staleness and requires external state management.
Sources. Kimi K3: Open Frontier Intelligence
Policy distillation
Intuition. Consolidate specialist behavior into one student.
Definition. Train a student policy from signals supplied by stronger or specialized teachers.
Formula or process. teacher -> student
Learn first. None
K3 adaptation. Consolidates nine experts.
Common misconception. Distillation need not average all teachers.
Boundary. Student capacity can bottleneck transfer.
Sources. Kimi K3: Open Frontier Intelligence
Reasoning effort
Intuition. Control how much computation a policy spends before finishing.
Definition. Train policy variants under problem-relative output budgets.
Formula or process. T(y) <= tau b0(x)
Learn first. reinforcement-learning
K3 adaptation. Uses low high and max experts.
Common misconception. Effort levels are not fixed published token counts.
Boundary. Token count is an imperfect compute proxy.
Sources. Kimi K3: Open Frontier Intelligence
Reference models
Intuition. Compare a learning policy against a fixed or lagged policy distribution.
Definition. Reference logits define regularization reward or divergence signals during post-training.
Formula or process. signal depends on log pi / pi_ref
Learn first. reinforcement-learning
K3 adaptation. Reference forwarding is streamed for long-context RL.
Common misconception. A reference model need not occupy one permanent full replica.
Boundary. Extra inference is expensive.
Sources. Kimi K3: Open Frontier Intelligence
Reinforcement learning
Intuition. Improve behavior using outcome-dependent feedback.
Definition. Update a policy to increase expected reward from sampled trajectories.
Formula or process. max E[r(y)]
Learn first. None
K3 adaptation. Trains domain and effort specialists.
Common misconception. Reward does not need a differentiable environment.
Boundary. Policies exploit reward defects.
Sources. Kimi K3: Open Frontier Intelligence
Supervised fine-tuning
Intuition. Learn desired behavior from verified demonstrations.
Definition. Optimize next-token likelihood on curated instruction trajectories.
Formula or process. min -log p(y|x)
Learn first. None
K3 adaptation. Establishes the agentic cold start.
Common misconception. SFT is not reinforcement learning.
Boundary. Demonstrations bound represented behavior.
Sources. Kimi K3: Open Frontier Intelligence
Task synthesis
Intuition. Generate new training problems from source materials and constraints.
Definition. A controlled pipeline converts concepts and retrieved evidence into tasks and evaluators.
Formula or process. concepts + materials -> task
Learn first. None
K3 adaptation. Uses a knowledge graph.
Common misconception. Fluent prompts need not be valid tasks.
Boundary. Verification is essential.
Sources. Kimi K3: Open Frontier Intelligence
Test-time scaling
Intuition. Spend more inference-time computation when a problem warrants it.
Definition. Vary reasoning tokens trajectories tools or verification effort after pretraining to improve task success.
Formula or process. quality=f(model, inference effort, verification)
Learn first. transformer-basics
K3 adaptation. Multi-effort RL and agentic environments train controllable effort policies.
Common misconception. A longer answer is not automatically better reasoning.
Boundary. Extra computation can amplify errors without reliable evaluation.
Sources. Kimi K3: Open Frontier Intelligence
Verifiable rewards
Intuition. Ground feedback in checkable outcomes.
Definition. Deterministic tests or independent evaluators map observable products and world states to reward.
Formula or process. r=verify(state,artifact)
Learn first. reinforcement-learning
K3 adaptation. Grounds search kernels agents and web tasks.
Common misconception. Self-reported completion is not verification.
Boundary. Verifiers can be incomplete or exploitable.
Sources. Kimi K3: Open Frontier Intelligence
White-box environments
Intuition. Own and reconfigure the agent harness during training.
Definition. Tools prompts context memory and protocols are explicit composable modules.
Formula or process. harness=compose(modules)
Learn first. agentic-environments
K3 adaptation. Varies harnesses by task group.
Common misconception. White-box does not mean trivial.
Boundary. Module diversity may miss real systems.
Sources. Kimi K3: Open Frontier Intelligence
Pretraining
Compute allocation
Intuition. Spend a fixed compute budget across model size, data, and experimental search.
Definition. Scaling experiments estimate configurations that minimize loss under a compute constraint.
Formula or process. C approximately 6ND
Learn first. scaling-laws
K3 adaptation. K3 jointly retunes batch, learning rate, tokens per parameter, and model shape.
Common misconception. The largest feasible model is not automatically compute-optimal.
Boundary. Allocation inherits uncertainty from fitted experiments.
Sources. Kimi K3: Open Frontier Intelligence
Data filtering
Intuition. Admit examples according to explicit quality and safety criteria.
Definition. Rule- and model-based classifiers accept, reject, or transform candidate training records.
Formula or process. keep(x)=1[quality(x) >= tau]
Learn first. data-curation
K3 adaptation. K3 reports domain-specific rules and classifiers.
Common misconception. A stricter threshold is not always a better dataset.
Boundary. Filters can suppress rare but valuable examples.
Sources. Kimi K3: Open Frontier Intelligence
Dataset deduplication
Intuition. Prevent repeated examples from silently dominating the learning distribution.
Definition. Exact and approximate similarity tests identify redundant records within and across sources.
Formula or process. reject(x) if similarity(x, seen) >= tau
Learn first. data-curation
K3 adaptation. K3 reports exact, fuzzy, perceptual, and structural checks by modality.
Common misconception. Near-duplicate detection is not exact hashing.
Boundary. Similarity thresholds trade contamination risk against false removal.
Sources. Kimi K3: Open Frontier Intelligence
Empirical scaling laws
Intuition. Fit how loss changes as model size, tokens, and compute change, then allocate a fixed budget.
Definition. A fitted empirical surface estimates held-out loss from controlled training runs and supports model/data/hyperparameter selection.
Formula or process. L(N,D)=L_inf + A N^{-alpha} + B D^{-beta}; C≈6ND
Learn first. None
K3 adaptation. K3 reports independent searches for cosine and WSD and an approximately 2.5x scaling-efficiency gain over K2.
Common misconception. A fitted curve is not a universal physical law and does not remove uncertainty.
Boundary. The repository surface is synthetic and does not reproduce paper coefficients.
Sources. Kimi K3: Open Frontier Intelligence
Language-model pretraining
Intuition. Learn general sequence structure by repeatedly predicting the next token.
Definition. Optimize a model over a broad token distribution before task-specific adaptation.
Formula or process. min_theta -sum_t log p_theta(x_t | x_<t)
Learn first. None
K3 adaptation. K3 applies one native next-token objective to interleaved text and vision tokens.
Common misconception. Pretraining data is not a neutral sample of the world.
Boundary. Objective quality depends on the data distribution and optimization system.
Sources. Kimi K3: Open Frontier Intelligence
Model–data tradeoffs
Intuition. More parameters and more tokens compete for the same finite compute.
Definition. Iso-compute comparisons vary model size and training-token count while holding approximate cost fixed.
Formula or process. D=C/(6N)
Learn first. compute-allocation
K3 adaptation. K3 reports improved scaling efficiency relative to K2.
Common misconception. Parameter count alone does not describe training investment.
Boundary. Tradeoffs vary across architectures, schedules, and datasets.
Sources. Kimi K3: Open Frontier Intelligence
Pretraining data curation
Intuition. Training data is an engineered distribution, not an unfiltered heap.
Definition. Source selection, quality filtering, deduplication, transformation, provenance, and mixture policy determine the examples presented to the optimizer.
Formula or process. accepted = quality(source) and not duplicate(source)
Learn first. None
K3 adaptation. K3 reports domain-specific filtering and sampling across Web Text, Code, Mathematics, Knowledge, and vision.
Common misconception. Removing more data is not automatically better; filters can erase rare valuable material.
Boundary. Local fixtures demonstrate decisions, not corpus quality at scale.
Sources. Kimi K3: Open Frontier Intelligence
Pretraining data mixtures
Intuition. Sampling weights decide how often each domain teaches the model.
Definition. A categorical source policy combines domain distributions into the training distribution.
Formula or process. P_train(x)=sum_d w_d P_d(x)
Learn first. data-curation
K3 adaptation. K3 tunes sampling rates with smaller ablation runs.
Common misconception. Mixture weights need not match raw corpus sizes.
Boundary. Optimal weights depend on model, budget, and evaluation priorities.
Sources. Kimi K3: Open Frontier Intelligence
Progressive long-context training
Intuition. Accepting a long tensor and learning to use distant evidence are different achievements.
Definition. A curriculum progressively increases sequence length while supplying cleaned long documents and tasks whose solution requires distant evidence.
Formula or process. 8K → 64K → 256K → 1M
K3 adaptation. NoPE avoids positional rescaling while KDA and sequence partitioning support progressive extension.
Common misconception. Context capacity alone does not demonstrate million-token retrieval or reasoning.
Boundary. Local chunking proves token/state mechanics, not semantic retention at one million tokens.
Sources. Kimi K3: Open Frontier Intelligence
Training recipe
Intuition. A large run is a synchronized control program, not merely an optimizer choice.
Definition. Optimizer, schedule, warmup, decay, clipping, precision, load balance, batch progression, and checkpoint policy jointly govern optimization.
Formula or process. lr(t)=warmup(t) then cosine_decay(t); weight_decay=0.1
Learn first. per-head-muon, quantile-balancing
K3 adaptation. K3 reports native multimodal next-token training, Per-Head Muon, weight clipping, QB, cosine decay, 1% warmup, and 0.1 weight decay.
Common misconception. Hyperparameters cannot be compared fairly across schedules without retuning them.
Boundary. Local schedule arithmetic is not a large-run stability result.
Sources. Kimi K3: Open Frontier Intelligence
Serving
Admission control
Intuition. Reject queue or defer work before it makes service guarantees impossible.
Definition. A policy checks predicted resource and latency cost against current capacity and budgets.
Formula or process. admit iff predicted demand <= safe capacity
Learn first. serving-scheduling
K3 adaptation. K3 reports budget-based fleet admission with cache awareness.
Common misconception. Admission is not identical to request routing.
Boundary. Prediction errors can waste capacity or violate budgets.
Sources. Kimi K3: Open Frontier Intelligence
Autoregressive decoding
Intuition. Extend a sequence one accepted token at a time.
Definition. Each step reads prior state emits a distribution selects tokens and updates cache.
Formula or process. state_t -> p(x_t+1), state_t+1
Learn first. model-serving
K3 adaptation. Specialized KDA state updates and speculative reconstruction serve K3.
Common misconception. Decode is not simply small-batch training.
Boundary. Launch and memory latency dominate small steps.
Sources. Kimi K3: Open Frontier Intelligence
Deployment optimization
Intuition. Shape training around serving constraints.
Definition. Incorporate precision memory and decoding requirements into model optimization.
Formula or process. quality subject to latency and memory
Learn first. None
K3 adaptation. Couples QAT and speculative draft training.
Common misconception. Deployment is not only a compiler problem.
Boundary. Gains depend on hardware and workload.
Sources. Kimi K3: Open Frontier Intelligence
EAGLE-3 draft model
Intuition. Train one cheap recurrent layer to propose several future tokens from multi-level target-model features.
Definition. A speculative draft fuses low middle and high target features and recurrently proposes tokens that the frozen target verifies losslessly.
Formula or process. target features -> recurrent draft proposals -> target verification
Learn first. multi-token-prediction, speculative-decoding
K3 adaptation. K3 initializes the draft from its pretrained MTP layer and unrolls it for seven steps.
Common misconception. The draft does not replace or approximate away final target verification.
Boundary. The repository does not reproduce acceptance rate or serving speedup.
Sources. Kimi K3: Open Frontier Intelligence, EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
Hybrid KDA-MLA prefix cache
Intuition. Reuse a prefix only when its token cache and recurrent state describe the same causal boundary.
Definition. Fine logical prefix hashes map to coarse physical pages carrying MLA KV while aligned KDA checkpoints restore recurrent state.
Formula or process. hit(b) iff MLA_KV[0:b] and KDA_state(b) coexist
Learn first. prefix-caching, latent-attention, recurrent-state
K3 adaptation. K3 packs both state types into one managed paged cache and reconstructs from aligned checkpoints.
Common misconception. Matching text with mismatched state boundaries is not a valid hit.
Boundary. The local miniature validates boundaries but not storage layout or performance.
Sources. Kimi K3: Open Frontier Intelligence
Model serving
Intuition. Turn model execution into a predictable multi-request service.
Definition. A serving system admits batches schedules executes caches and returns inference requests under latency and capacity constraints.
Formula or process. service = admission + scheduling + execution + state
Learn first. distributed-systems
K3 adaptation. Manages hybrid KDA and MLA state across a fleet.
Common misconception. A fast kernel alone is not a serving system.
Boundary. Workload mix determines performance.
Sources. Kimi K3: Open Frontier Intelligence
Prefill
Intuition. Process an existing prompt before incremental generation.
Definition. Prefill computes model states and caches for all prompt tokens.
Formula or process. prompt[1:T] -> state_T + cache
Learn first. model-serving
K3 adaptation. Long KDA prefill uses intra-device sequence parallel work.
Common misconception. Prefill and decoding have different shapes.
Boundary. Long prompts can dominate time to first token.
Sources. Kimi K3: Open Frontier Intelligence
Prefix caching
Intuition. Reuse model state produced by an identical prompt prefix.
Definition. Hash-addressed cached state avoids recomputing a matched causal prefix.
Formula or process. hash(prefix) -> reusable state at boundary b
Learn first. model-serving
K3 adaptation. K3 jointly restores MLA KV and KDA checkpoints.
Common misconception. A text match is insufficient if state boundaries disagree.
Boundary. Memory admission and invalidation constrain hit rate.
Sources. Kimi K3: Open Frontier Intelligence
Serving scheduling
Intuition. Decide where and when inference requests run.
Definition. A scheduler assigns admitted requests to replicas and batches under latency capacity and fairness constraints.
Formula or process. request -> feasible replica and time slot
Learn first. model-serving
K3 adaptation. K3 combines cache affinity with budget-aware admission.
Common misconception. Cache locality is not the only objective.
Boundary. Future request cost is uncertain.
Sources. Kimi K3: Open Frontier Intelligence
Speculative decoding
Intuition. Let a cheap draft propose tokens that a target verifies.
Definition. Lossless sampling accepts shared probability mass and corrects rejected proposals.
Formula or process. A=sum min(p,q)
Learn first. None
K3 adaptation. Uses an EAGLE-3-style draft.
Common misconception. Argmax equality is not probabilistic verification.
Boundary. Speedup depends on acceptance and verification cost.
Sources. Kimi K3: Open Frontier Intelligence
Sparse Capacity
Expert load balancing
Intuition. Prevent a few experts or ranks from becoming the throughput bottleneck.
Definition. A routing control targets a more even assignment distribution while preserving useful specialization.
Formula or process. load_j approx tokens*k/experts
Learn first. sparse-routing
K3 adaptation. Quantile Balancing adjusts selection thresholds without an auxiliary loss.
Common misconception. Equal token counts do not guarantee equal rank time.
Boundary. Token cost and communication can remain uneven.
Sources. Switch Transformers
Gated linear units
Intuition. Let one learned branch control how much of another branch passes.
Definition. Two projections form a gate and content branch whose element-wise product is transformed downstream.
Formula or process. GLU(x)=gate(xW_g) odot xW_u
Learn first. activation-functions
K3 adaptation. SiTU-GLU soft-caps both routed-expert branches.
Common misconception. Gating does not itself bound either branch.
Boundary. Extra projections increase compute and parameters.
Sources. Language Modeling with Gated Convolutional Networks
Quantile Balancing
Intuition. Move each expert's selection threshold to target a desired load without changing mixture weights.
Definition. Expert-specific routing biases are derived from score-margin quantiles and frozen for inference.
Formula or process. b_j <- quantile(score margin for expert j)
Learn first. moe-routing
K3 adaptation. Balances 896 routed experts without an auxiliary balancing loss.
Common misconception. Selection bias is not added to the final mixture weight.
Boundary. Small simulations do not prove cluster-scale balance.
Sources. Kimi K3: Open Frontier Intelligence
SiTU-GLU
Intuition. Preserve gated expressivity while softly capping both branches.
Definition. Sigmoid–tanh gate and tanh up branch constrain routed-expert activation magnitude.
Formula or process. (beta_1 tanh(x_g/beta_1) sigmoid(x_g)) odot (beta_2 tanh(x_u/beta_2))
Learn first. swiglu
K3 adaptation. Stabilizes extreme-sparsity LatentMoE computation.
Common misconception. Bounded activations do not solve routing load imbalance.
Boundary. The miniature demonstrates the bound, not trillion-parameter training stability.
Sources. Kimi K3: Open Frontier Intelligence
Sparse routing
Intuition. Send each token to only a small subset of available experts.
Definition. A router scores experts selects a subset and combines their outputs using normalized weights.
Formula or process. y=sum_{j in selected(x)} p_j E_j(x)
Learn first. mixture-of-experts
K3 adaptation. K3 selects 16 of 896 routed experts alongside a shared path.
Common misconception. Sparse arithmetic does not guarantee balanced device work.
Boundary. Routing introduces dispatch communication and capacity constraints.
Sources. Switch Transformers
SwiGLU
Intuition. Multiply a smooth learned gate by a learned value branch.
Definition. A Swish-activated gate modulates a parallel linear projection.
Formula or process. SwiGLU(x)=Swish(xW_g) odot (xW_u)
Learn first. None
K3 adaptation. K3 replaces unbounded branches with SiTU-GLU in routed experts.
Common misconception. Gating does not impose a hard activation bound.
Boundary. Numerical behavior depends on inputs, weights, precision, and surrounding normalization.
Sources. GLU Variants Improve Transformer
Swish activation
Intuition. Smoothly gate a value by its sigmoid.
Definition. Swish multiplies an input by its logistic sigmoid and is used in SwiGLU gates.
Formula or process. swish(x)=x sigmoid(x)
Learn first. activation-functions
K3 adaptation. SiTU replaces the unbounded Swish factor with a tanh-soft-capped gate branch.
Common misconception. Swish is smooth but unbounded above.
Boundary. Numerical range still follows input scale.
Sources. GLU Variants Improve Transformer
Top-k mixture-of-experts routing
Intuition. Score many specialists but execute only a small subset for each token.
Definition. A router selects the highest-scoring experts and combines their outputs with normalized routing weights.
Formula or process. r=TopK(s); y=sum_{i in r} p_i E_i(x)
Learn first. None
K3 adaptation. Stable LatentMoE scores 896 routed experts and activates 16 per token.
Common misconception. Sparse activation does not mean the model contains only the active parameters.
Boundary. Routing quality and systems balance are separate problems.
Sources. Switch Transformers
Top-k routing
Intuition. Choose the k highest-scoring experts for each token.
Definition. Expert logits are ranked and only the largest k participate in routed computation.
Formula or process. selected(x)=TopK(router(x),k)
Learn first. sparse-routing
K3 adaptation. Stable LatentMoE uses top-16 paper routing while the miniature uses top-2.
Common misconception. Selection bias need not equal mixture weight.
Boundary. Discrete boundaries can amplify small score shifts.
Sources. Switch Transformers