Loss Function
损失函数EssentialA single number measuring how far a model's prediction is from the correct answer; training just means shrinking it.
A loss function measures how badly wrong a single prediction is, producing a real number that grows as the gap between prediction and target grows. During training, backpropagation computes the gradient of the loss with respect to every parameter, and the optimizer updates the parameters in the direction that shrinks the loss, repeating this until the loss stops improving meaningfully. Regression over continuous values commonly uses mean squared error (MSE) or L1 loss; classification, or predicting the next token, commonly uses cross-entropy. “Objective function” is a broader term — it can be a loss to minimize, or a quantity to maximize instead, such as cumulative reward in reinforcement learning. In embodied models, the choice of loss directly affects action quality: ACT uses L1 loss to regress actions, while diffusion policies and flow-matching models use a denoising loss or a flow-matching loss instead.
ExampleThe ACT paper found that L1 loss modeled the action sequence more precisely than the more common L2 (mean squared error) loss when regressing actions, so they switched to L1.
- Also called
- Cost Function, Objective Function
- Related
- Mean Squared Error · L1 Loss · Cross-Entropy · Denoising Loss (Diffusion Loss) · Gradient Descent · Backpropagation
- Sources
- Google Machine Learning Glossary: loss function
Wikipedia: Loss function
Zhao et al. 2023: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)