WORLD 04 INFERENCE ENGINEERING
Watch every token earn its place.
The AI Inference Architecture Handbook, animated. From silicon and quantization to serving engines, distributed scale and FinOps — 24 scenes across 8 modules.
MODULE 01 · 3 SCENES
Foundations of Inference & Execution
Computational traits of the forward pass, the end-to-end request lifecycle, and the metric split that defines every serving SLO.
01 · FOUNDATIONS
One token at a time
An LLM never writes a sentence — it predicts the next token, then does it again.
OPEN →02 · TWO PHASES
Prefill reads. Decode writes.
Your prompt is ingested in parallel; generation streams one token per step.
OPEN →03 · METRICS & SLOs
TTFT, TPOT and goodput
The three numbers every inference SLO is made of.
OPEN →MODULE 02 · 3 SCENES
Hardware Micro-Architecture & Silicon
CPUs to LPUs, the Von Neumann bottleneck, warp scheduling on the H100, and the edge-versus-datacenter split.
04 · SILICON TOPOLOGY
CPU, GPU, TPU, LPU
Every substrate makes a different bet on parallelism, memory and flexibility.
OPEN →05 · MEMORY BANDWIDTH
The wall every model hits
Why decode stops being compute-bound and starts waiting on bytes.
OPEN →06 · GPU MICRO-ARCH
Warps and Tensor Cores
Threads, SMs and matrix engines — how software maps onto H100 silicon.
OPEN →MODULE 03 · 3 SCENES
Transformation, Compilation & Quantization
Precision ladders and PTQ math, graph fusion and pruning, and the compilers that lower model code to bare metal.
07 · QUANTIZATION
The precision ladder
FP32 down to FP4: what each rung saves, and what calibration rescues.
OPEN →08 · GRAPH OPTIMIZATION
Fuse the graph
Kernel fusion, dead-node elimination and constant folding shrink the launch tax.
OPEN →09 · AI COMPILERS
Model code to bare metal
TVM, OpenXLA, TensorRT and MAX lower high-level graphs to tuned kernels.
OPEN →MODULE 04 · 3 SCENES
The Runtime & Memory Management Layer
Engine landscape, the KV cache problem, PagedAttention, radix-tree prefix reuse, and kernel-level attention.
10 · THE KV ECONOMY
Half a megabyte per token
Attention remembers its past in the KV cache — memory becomes the budget.
OPEN →11 · PAGED ATTENTION
End the fragmentation tax
Reserve exact blocks on demand instead of worst-case contiguous slabs.
OPEN →12 · PREFIX CACHING
Reuse what shared prompts pay for
Radix trees match cached prefixes so multi-turn history is computed once.
OPEN →MODULE 05 · 3 SCENES
Interaction, Sampling & Constrained Generation
The autoregressive token loop, distribution filters, and grammar-masked structured outputs.
13 · THE TOKEN LOOP
Logits in, token out
Temperature, top-k and top-p sculpt probabilities before the dice roll.
OPEN →14 · SAMPLING FILTERS
Shape the distribution
Low-level logic of the filters — and what each costs on hardware.
OPEN →15 · STRUCTURED OUTPUTS
Force valid JSON
Grammar constraints mask impossible tokens before the softmax ever runs.
OPEN →MODULE 06 · 4 SCENES
Serving Middleware, Frameworks & Scheduling
vLLM and SGLang internals, continuous batching, chunked prefill, PD disaggregation, and speculative decoding.
16 · BATCHING
Never leave a slot empty
A request finishes? Swap the next one in mid-generation.
OPEN →17 · CHUNKED PREFILL
Decode never freezes
Slice big prompts into micro-chunks so active decodes keep ticking.
OPEN →18 · PD DISAGGREGATION
Two pools, two appetites
Compute-hungry prefill and memory-bound decode run on separate GPUs.
OPEN →19 · SPECULATIVE DECODING
Draft cheaply, verify once
A small model guesses several tokens; the big model checks them in one pass.
OPEN →MODULE 07 · 3 SCENES
Distributed Inference & Infrastructure Scale
Tensor, pipeline and expert parallelism, KV cache distribution, and edge-cloud hybrid pools.
20 · PARALLELISM
Slice weights, chain layers
Tensor parallelism inside a node, pipeline parallelism across it.
OPEN →21 · EXPERT PARALLELISM
Route to a few experts
MoE keeps trillion-parameter capacity while activating a sliver per token.
OPEN →22 · DISTRIBUTED CACHE
KV caches across nodes
Offload, stream and balance attention memory over the network fabric.
OPEN →MODULE 08 · 2 SCENES
Production Operations, Deployment & FinOps
Containers and KServe, cold starts, observability of saturation and drift, and run-cost optimization.