Embodied AI Glossary中文

Entropy Regularization

熵正则化Advanced

Adding a bonus for policy entropy to the training objective, to encourage the policy to stay random and keep exploring.

Entropy measures how random a probability distribution is: a uniform distribution has high entropy, and a nearly certain one has low entropy. Entropy regularization adds α times the policy's entropy to the reinforcement-learning objective (α is a weighting coefficient), so the agent pursues return while retaining some randomness, avoiding collapsing onto a suboptimal behavior too early. There are two common forms: adding a small entropy bonus term to the loss in policy-gradient algorithms like PPO, or writing entropy directly into the optimization objective, giving maximum-entropy reinforcement learning — the leading example is Haarnoja and colleagues' 2018 SAC (Soft Actor-Critic), whose Q-value target also carries an entropy term. Too large an α and the policy stays erratic forever; too small and it risks entropy collapse, stopping exploration too soon.

Examplelegged_gym's default PPO config sets the entropy coefficient entropy_coef to 0.01, adding 0.01 times the policy's entropy as a bonus in the loss.

Also called
Entropy Bonus, Maximum Entropy RL
Related
Soft Actor-Critic · Proximal Policy Optimization · Exploration vs. Exploitation · Entropy Collapse / Mode Collapse · KL Regularization · Deterministic vs. Stochastic Policy
Sources
OpenAI Spinning Up: Soft Actor-Critic (Entropy-Regularized RL)
Soft Actor-Critic: Off-Policy Maximum Entropy Deep RL with a Stochastic Actor (arXiv:1801.01290)
legged_gym: legged_robot_config.py

See it in the full glossary →