Gradient Clipping
梯度裁剪AdvancedScaling down an overly large gradient to stay within a threshold, so one bad update can't wreck the model.
Training occasionally hits exploding gradients: one step's gradient is unusually large, pushing the parameters much too far and sending the loss spiking or even to NaN. Gradient clipping checks the gradient's magnitude before the optimizer applies it, and shrinks it if it exceeds a threshold. The most common form is norm clipping: treat every parameter's gradient as one long vector, compute its total norm, and if that exceeds max_norm, multiply the whole vector by max_norm divided by the total norm, which keeps the direction the same and only shortens its length; an alternative, value clipping, clamps each individual component to a fixed range. Pascanu, Mikolov, and Bengio proposed gradient-norm clipping in 2013 while analyzing exploding gradients in recurrent neural networks. It's now essentially a default setting when training Transformers and diffusion policies, and common PPO implementations generally include it too.
ExampleIn PyTorch, calling torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) after loss.backward() and before optimizer.step().
- Also called
- Gradient Norm Clipping
- Related
- Vanishing / Exploding Gradients · Learning Rate · Optimizer · Backpropagation · Recurrent Neural Network · Proximal Policy Optimization
- Sources
- On the difficulty of training Recurrent Neural Networks (arXiv 1211.5063)
torch.nn.utils.clip_grad_norm_ (PyTorch 文档) (Chinese)