AI ENGVisual Encyclopedia

SCENE 06 / 24 · HARDWARE MICRO-ARCHITECTURE & SILICON

Warps and Tensor Cores

Threads, SMs and matrix engines — how software maps onto H100 silicon.

RESIDENT WARPS · 1,690 / 8448SCHEDULER BUBBLE · 80%
SM · teal = warps residentgold = Tensor Core MMA active132 SMs · 64 warp slots · 4 TC each

RUN to sweep occupancy from starved to saturated.

TECHNICAL BREAKDOWNModule 2: Hardware Micro-Architecture & Silicon Substrates

Warps, SMs and Tensor Cores: mapping software onto H100

A CUDA thread is a fiction the scheduler resolves into 32-thread warps. Each of the H100's 132 SMs hosts up to 64 resident warps; when one stalls on memory, another issues. Tensor Cores are the matrix engines inside — fused multiply-accumulate tiles that give the GPU its headline FLOPs.

Latency Hiding by Occupancy

Warp schedulers swap in ready warps every cycle. Resident warp density is the currency: too few warps and every HBM stall becomes dead silicon.

Tensor Cores

Each SM packs 4 Tensor Cores executing 16×8×16 MMA tiles. Peak FLOPs live here; a kernel that doesn't feed them (via WMMA/MMA fragments) leaves 10× on the table.

Registers Are the Real Cache

Fused kernels keep intermediates in registers and shared memory; unfused chains pay HBM round-trips per op. This is why fusion matters more than kernel micro-tuning.

MATHEMATICAL FORMULATION · OCCUPANCY BUBBLE
bubble ≈ 1 − resident_warps / (SMs × warps_per_SM) · H100: 132 × 64 = 8,448 warp slots

At 30% occupancy, ~5,000 warp slots idle. Latency bubbles appear exactly where resident density drops — the visual signature of an under-occupied kernel.

REAL-WORLD PRODUCTION ENGINEERING
  • FlashAttention's insight is scheduling, not math: tile attention so the softmax chain never leaves SRAM.
  • Nsight Compute's 'achieved occupancy' vs 'theoretical occupancy' gap is the first thing kernel engineers check.