AdamW Mechanics, Decoupled Weight Decay, & Cosine Schedules
Training stability and convergence speed rely heavily on optimizer mechanics. AdamW decouples weight decay from gradient updates, paired with linear warmup and cosine decay schedules.
AdamW vs. L2 Regularization
In standard Adam, L2 weight decay is added to gradients, interacting undesirably with moving variance averages. AdamW applies weight decay directly to weight parameters (w = w - lr * wd * w), dramatically improving deep LLM stability.
Warmup & Cosine Decay
1. Warmup: LR ramps linearly from 0 to max LR over the first 1-2% of steps to prevent gradient explosion during volatile early initialization. 2. Cosine Decay: LR decays along a cosine curve down to 10% peak LR at run completion.
Gradient Clipping
If global gradient L2 norm exceeds a threshold (typically ||g|| > 1.0), all gradients are rescaled: g = g * (threshold / ||g||) to suppress extreme loss spikes.
Smoothly reducing learning rate from peak target down to minimum floor.
- Llama 3 used a peak learning rate of 3.0e-4 for 8B and 1.5e-4 for 70B, with AdamW hyper-parameters β1=0.9, β2=0.95.