Learning Rate Schedule (Warmup and Cosine Decay)
学习率调度(预热与余弦退火)CommonChanging the learning rate over the course of training: ramping it up during warmup, then easing it down along a cosine curve.
A learning rate schedule changes the learning rate during training according to a predefined rule, rather than keeping it fixed throughout. The most common combination is “warmup plus cosine decay.” Warmup linearly raises the learning rate from near 0 up to its peak value over the first stretch of training: at that point the parameters are freshly initialized or a new module has just been attached, so gradients are unstable, and jumping straight to a large learning rate risks divergence. Goyal and colleagues (2017) used gradual warmup together with a linear scaling rule to scale ResNet-50 training up to a batch size of 8192 without losing accuracy. After warmup, the learning rate eases down to a very small value along a cosine curve, letting the model converge carefully in the later stages; this cosine shape comes from Loshchilov and Hutter's SGDR paper (ICLR 2017). Other common schedules include step decay, exponential decay, and linear decay.
Exampleopenpi's default recipe for fine-tuning π0 linearly ramps the learning rate up to 2.5e-5 over the first 1,000 steps, then eases it down to 2.5e-6 along a cosine curve over 30,000 steps.
- Also called
- Warmup, Cosine Decay, Cosine Annealing, Learning Rate Decay
- Related
- Learning Rate · Optimizer · AdamW · Batch Size · Convergence · Fine-tuning
- Sources
- SGDR: Stochastic Gradient Descent with Warm Restarts (arXiv 1608.03983)
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour (arXiv 1706.02677)
openpi optimizer.py(CosineDecaySchedule 默认参数) (Chinese)