AI ENGVisual Encyclopedia

SCENE 05 / 24 · HARDWARE MICRO-ARCHITECTURE & SILICON

The wall every model hits

Why decode stops being compute-bound and starts waiting on bytes.

ARITHMETIC INTENSITY · 1000 FLOP/BPREFILL REGIME (COMPUTE-BOUND)
ARITHMETIC INTENSITY (FLOP per byte, log scale) →TFLOPS0.11101001000GPU Peak (989 TF)LPU Peak (750 TF)Knee 295 FLOP/BKnee 9.4 FLOP/BGPU 989 TFLPU 750 TF3.35 TB/s HBM80 TB/s SRAM

Prefill regime (high intensity): both accelerators hit their compute ceilings — the GPU's 989 TF peak leads.

TECHNICAL BREAKDOWNModule 2: Hardware Micro-Architecture & Silicon Substrates

The roofline: why decode lives on the memory wall

Every operation has an arithmetic intensity — FLOPs per byte moved. Above the knee, you're compute-bound; below it, you're streaming bytes and the bandwidth roof caps you. Inference's two phases sit on opposite sides, and almost every inference optimization is an attempt to move work up the intensity axis.

Von Neumann in Production

Data shuttles between memory and compute over a bus narrower than either side. SRAM (fast, tiny) vs HBM (big, slow) defines the trade every accelerator makes.

Decode's Math

One 7B FP16 token = 14 GB of weights, ~2×7 GFLOPs of compute. Intensity ≈ 1 FLOP/byte. Peak FLOPs are irrelevant; the bandwidth roof is destiny.

Escalation Ladders

Batching shares weight loads across requests; quantization shrinks the bytes; fusion keeps intermediates on-chip. All three are intensity upgrades.

MATHEMATICAL FORMULATION · ROOFLINE
achievable_TFLOPs = min(peak, BW(TB/s) × AI(FLOP/B)) · AI_decode = 2P / (P × b) = 2/b

For FP16 (b=2 bytes/param) decode AI ≈ 1. The H100 compute roof (989 TF) is 300× above its bandwidth roof at that intensity (3.35 TF). Utilization is not a bug — it is geometry.

REAL-WORLD PRODUCTION ENGINEERING
  • Batch size is the cheapest intensity knob: 64-way batching multiplies AI by ~64 without touching the model.
  • NVLink-class interconnects exist because tensor parallelism multiplies bytes moved per FLOP even further.