Convergence
收敛CommonThe point in training where the loss or return stops changing much, meaning the model has settled into a stable state.
Convergence describes training reaching a point where the loss (a number measuring the gap between prediction and target) or, in reinforcement learning, the return, stops changing much across iterations, so further training gives little benefit. Gradient-descent-style methods are only guaranteed to converge to a local optimum or stationary point under certain conditions; in practice, deep networks are judged by watching the training curve. Too large a learning rate can make the loss oscillate or even diverge, and exploding gradients or data problems can also prevent convergence. In robot learning, a converged loss doesn't necessarily mean the best policy: large-scale robomimic experiments found that the training objective and the real evaluation objective don't always agree, so picking a checkpoint by validation loss alone is unreliable — it's usually necessary to actually roll out the policy and check its success rate before choosing a model.
ExampleWhen training a legged locomotion policy, the average-return curve rises quickly and then flattens out, which is taken as roughly converged. When training a behavior-cloning policy, even after the loss curve flattens, several checkpoints still get evaluated for success rate before deciding which one to use.
- Also called
- Training Convergence, Converge
- Related
- Loss Function · Learning Rate · Gradient Descent · Overfitting · Checkpoint · Early Stopping
- Sources
- Google Machine Learning Glossary: convergence
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (robomimic, arXiv 2108.03298)