Learning Rate
学习率LRCommonThe size of the step taken at each parameter update in gradient descent; one of the most important hyperparameters.
The learning rate is the hyperparameter in an optimization algorithm that controls how big each parameter update is: parameters move in the direction opposite the gradient, by an amount equal to the learning rate times the gradient (adaptive optimizers like Adam further rescale this per parameter, based on that parameter's gradient history). Set it too high and parameters overshoot the minimum, making the loss oscillate or even diverge to NaN; set it too low and convergence is slow, or training gets stuck at a mediocre point. It's usually the first thing tuned, and it's tied to batch size — when batch size goes up, the learning rate is often scaled up proportionally too (the linear scaling rule). In practice, training rarely uses one fixed value throughout; instead it follows a learning rate schedule, typically warming up first and then decaying gradually. Fine-tuning a pretrained large model usually uses a much smaller learning rate than training from scratch, to avoid damaging what it already learned.
ExampleLLaVA uses a learning rate of 2e-3 in its first stage, when only the projection layer is trained, and drops to 2e-5 in the second stage, when the language model is fine-tuned along with it. OpenVLA trains with a fixed 2e-5 throughout, and its authors found adding warmup gave no benefit; ACT uses 1e-5.
- Also called
- LR, Step Size
- Related
- Learning Rate Schedule (Warmup and Cosine Decay) · Gradient Descent · Optimizer · AdamW · Batch Size · Hyperparameter
- Sources
- Learning rate(Wikipedia)
Visual Instruction Tuning (LLaVA, arXiv 2304.08485)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)