AI ENGVisual Encyclopedia

SCENE 02 / 24 · FOUNDATIONS OF INFERENCE & EXECUTION

Prefill reads. Decode writes.

Your prompt is ingested in parallel; generation streams one token per step.

READY
COMPUTE UNITS28%
MEMORY BANDWIDTH92%

Two phases live inside every request. RUN to feel the handoff.

TECHNICAL BREAKDOWNModule 1: Foundations of Inference & Execution

Two personalities of one GPU: compute-bound prefill, bandwidth-bound decode

Every request has two personalities. Prefill ingests the whole prompt in one parallel sweep — thousands of tokens flow through every layer together, saturating tensor cores. Decode emits one token at a time, streaming the full weight matrix from HBM for each: compute idles while memory works. The handoff between these regimes drives most serving architecture.

Prefill = FLOP Party

With B×N tokens in flight, matmuls run near peak TFLOPS. Arithmetic intensity is high; the compute roof of the roofline model is what limits you.

Decode Is a Memory Stream

Batch of 1 decode moves ~14 GB of weights to produce one token: intensity near 1 FLOP/byte. The GPU spends its time waiting on HBM, at maybe 1-5% of peak FLOPs.

Why Batching Works

Amortization: 64 concurrent decodes read the same weights once per step. Batched decode pushes intensity back toward the compute roof — this is the economic engine of every serving stack.

MATHEMATICAL FORMULATION · THE TWO ROOFS
AI = FLOPs / Bytes_moved · achievable ≈ min(peak_FLOPs, BW × AI)

Prefill of 2k tokens: intensity in the hundreds, compute-bound. Decode: intensity ≈ 1-2, bandwidth-bound. Same GPU, two ceilings — which is why prefill/decode scheduling and disaggregation exist.

REAL-WORLD PRODUCTION ENGINEERING
  • TTFT SLOs are prefill budgets; TPOT SLOs are bandwidth budgets. Marketing quotes one, users feel both.
  • vLLM logs prefill and decode as separate phases precisely because capacity planning treats them as different workloads.