Embodied AI Glossary中文

Gradient Descent

梯度下降Common

An optimization method that nudges parameters step by step opposite the loss function's gradient, making the loss smaller each time.

Gradient descent is the core optimization method for training neural networks, generally credited to the mathematician Augustin-Louis Cauchy in 1847. The gradient is the loss function's partial derivative with respect to each parameter, pointing in the direction the loss increases fastest; each step moves the parameters a bit in the opposite direction, and how far is set by the learning rate — too small and convergence is slow, too large and the loss oscillates or even diverges. Because deep-learning datasets are large, each step usually estimates the gradient from just a small batch of data, called stochastic (mini-batch) gradient descent, or SGD. The gradient itself is computed by backpropagation. The Adam and AdamW optimizers commonly used to train large models like VLAs today are refinements of gradient descent that add momentum and adaptive step sizes.

ExampleTraining a behavior-cloning policy: each step takes 64 frames of demonstration, computes the mean squared error between predicted and demonstrated actions, runs backpropagation to get the gradient, then updates once via “parameter ← parameter − learning rate × gradient.”

Also called
Stochastic Gradient Descent, SGD, Mini-batch Gradient Descent
Related
Backpropagation · Learning Rate · Optimizer · Loss Function · Batch Size · AdamW
Sources
Wikipedia: Gradient descent
Google Machine Learning Glossary

See it in the full glossary →