Embodied AI Glossary中文

Optimizer

优化器Common

The algorithm that turns gradients into parameter updates during training, such as SGD, Adam, or AdamW.

When training a neural network, backpropagation computes the loss's gradient with respect to every parameter, and the optimizer decides how to turn that gradient into an actual update: which direction to move, and how big a step to take. The most basic is stochastic gradient descent (SGD), smoother once momentum is added; Adam (Kingma and Ba, 2014) tracks both the mean and the squared mean of the gradient, giving each parameter its own adaptive step size, and needs little tuning while converging fast; AdamW (Loshchilov and Hutter, ICLR 2019) pulls weight decay out of the gradient update, generalizing better, and is common when training Transformers, diffusion policies, and similar models. Optimizers are usually configured together with a learning rate schedule (warmup, cosine decay) and gradient clipping. Note that PNDbotics' robot named “Adam” is an unrelated humanoid robot product, not the Adam optimizer.

ExampleDiffusion Policy's official training config uses torch.optim.AdamW with a learning rate of 1e-4, weight decay of 1e-6, and 500 warmup steps followed by cosine decay; in PyTorch, each training step calls zero_grad(), loss.backward(), and optimizer.step() in turn.

Also called
Optimization Algorithm
Related
Gradient Descent · Backpropagation · Learning Rate · AdamW · Learning Rate Schedule (Warmup and Cosine Decay) · Gradient Clipping
Sources
Adam: A Method for Stochastic Optimization
Decoupled Weight Decay Regularization (AdamW)
PyTorch Documentation: torch.optim

See it in the full glossary →