Two personalities of one GPU: compute-bound prefill, bandwidth-bound decode
Every request has two personalities. Prefill ingests the whole prompt in one parallel sweep — thousands of tokens flow through every layer together, saturating tensor cores. Decode emits one token at a time, streaming the full weight matrix from HBM for each: compute idles while memory works. The handoff between these regimes drives most serving architecture.
Prefill = FLOP Party
With B×N tokens in flight, matmuls run near peak TFLOPS. Arithmetic intensity is high; the compute roof of the roofline model is what limits you.
Decode Is a Memory Stream
Batch of 1 decode moves ~14 GB of weights to produce one token: intensity near 1 FLOP/byte. The GPU spends its time waiting on HBM, at maybe 1-5% of peak FLOPs.
Why Batching Works
Amortization: 64 concurrent decodes read the same weights once per step. Batched decode pushes intensity back toward the compute roof — this is the economic engine of every serving stack.
Prefill of 2k tokens: intensity in the hundreds, compute-bound. Decode: intensity ≈ 1-2, bandwidth-bound. Same GPU, two ceilings — which is why prefill/decode scheduling and disaggregation exist.
- TTFT SLOs are prefill budgets; TPOT SLOs are bandwidth budgets. Marketing quotes one, users feel both.
- vLLM logs prefill and decode as separate phases precisely because capacity planning treats them as different workloads.