AI ENGVisual Encyclopedia

MODULE 6: EXECUTION & DIAGNOSTICS · SCENE 20

AdamW and the cosine slide

Decoupled weight decay, bias-corrected moments, warmup, and a learning rate that coasts downhill.

STEP 200 / 1000LR MULTIPLIER: 0.927x

Warmup ramps learning rate from 0 during early steps; cosine decay smoothly coaxes parameters toward the local minimum.

TECHNICAL BREAKDOWNModule 6: Execution Dynamics & Diagnostics

AdamW Mechanics, Decoupled Weight Decay, & Cosine Schedules

Training stability and convergence speed rely heavily on optimizer mechanics. AdamW decouples weight decay from gradient updates, paired with linear warmup and cosine decay schedules.

AdamW vs. L2 Regularization

In standard Adam, L2 weight decay is added to gradients, interacting undesirably with moving variance averages. AdamW applies weight decay directly to weight parameters (w = w - lr * wd * w), dramatically improving deep LLM stability.

Warmup & Cosine Decay

1. Warmup: LR ramps linearly from 0 to max LR over the first 1-2% of steps to prevent gradient explosion during volatile early initialization. 2. Cosine Decay: LR decays along a cosine curve down to 10% peak LR at run completion.

Gradient Clipping

If global gradient L2 norm exceeds a threshold (typically ||g|| > 1.0), all gradients are rescaled: g = g * (threshold / ||g||) to suppress extreme loss spikes.

MATHEMATICAL FORMULATION · COSINE LEARNING RATE SCHEDULE
η_t = η_min + 0.5 × ( η_max - η_min ) × ( 1 + cos( π × ( t - t_warmup ) / ( T_total - t_warmup ) ) )

Smoothly reducing learning rate from peak target down to minimum floor.

REAL-WORLD PRODUCTION ENGINEERING
  • Llama 3 used a peak learning rate of 3.0e-4 for 8B and 1.5e-4 for 70B, with AdamW hyper-parameters β1=0.9, β2=0.95.