Embodied AI Glossary中文

Regularization

正则化Common

Adding constraints or penalties during training to keep a model from memorizing training data, improving performance on new data.

Regularization is a broad category of techniques for preventing overfitting: whenever a model keeps improving on the training set but gets worse on new data, regularization is used to limit the model's effective complexity. Explicit regularization adds a penalty term to the loss function, such as L2 regularization (also called weight decay, which penalizes the sum of squared weights and shrinks them overall) and L1 regularization (which penalizes absolute value and pushes many weights to exactly zero); implicit regularization includes early stopping, dropout (randomly zeroing some neurons during training), and data augmentation. Reinforcement learning has several regularizers of its own: entropy regularization encourages the policy to stay somewhat random to keep exploring, KL regularization keeps a fine-tuned policy from drifting too far from the original model, and behavior regularization keeps an offline reinforcement-learning policy close to the actions in the dataset.

ExampleTraining a visuomotor policy applies random cropping to camera images (data augmentation) and sets weight decay in the AdamW optimizer; RLHF subtracts a KL penalty term from the reward, keeping the model from drifting away from the original language model just to rack up a higher score.

Also called
Regularization Term
Related
Overfitting · Dropout · Early Stopping · Data Augmentation · Entropy Regularization · KL Regularization
Sources
Wikipedia: Regularization (mathematics)
Dive into Deep Learning: Weight Decay

See it in the full glossary →