Embodied AI Glossary中文

Gradient Clipping

梯度裁剪Advanced

Scaling down an overly large gradient to stay within a threshold, so one bad update can't wreck the model.

Training occasionally hits exploding gradients: one step's gradient is unusually large, pushing the parameters much too far and sending the loss spiking or even to NaN. Gradient clipping checks the gradient's magnitude before the optimizer applies it, and shrinks it if it exceeds a threshold. The most common form is norm clipping: treat every parameter's gradient as one long vector, compute its total norm, and if that exceeds max_norm, multiply the whole vector by max_norm divided by the total norm, which keeps the direction the same and only shortens its length; an alternative, value clipping, clamps each individual component to a fixed range. Pascanu, Mikolov, and Bengio proposed gradient-norm clipping in 2013 while analyzing exploding gradients in recurrent neural networks. It's now essentially a default setting when training Transformers and diffusion policies, and common PPO implementations generally include it too.

ExampleIn PyTorch, calling torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) after loss.backward() and before optimizer.step().

Also called
Gradient Norm Clipping
Related
Vanishing / Exploding Gradients · Learning Rate · Optimizer · Backpropagation · Recurrent Neural Network · Proximal Policy Optimization
Sources
On the difficulty of training Recurrent Neural Networks (arXiv 1211.5063)
torch.nn.utils.clip_grad_norm_ (PyTorch 文档) (Chinese)

See it in the full glossary →