Embodied AI Glossary中文

Entropy Collapse / Mode Collapse

熵坍缩 / 模式坍缩Advanced

A model's output diversity collapsing, so it only ever produces a handful of answers or actions.

These two terms describe the same broad phenomenon. Mode collapse was first used for GANs (generative adversarial networks): Goodfellow's 2016 tutorial describes it as the generator mapping many different input noise vectors to the same output, covering only one or two “modes” of the data distribution. Entropy collapse is used more in reinforcement learning: a policy's entropy (how random it is) drops rapidly early in training, the model settles on almost one single behavior, stops exploring, and performance plateaus; Cui and colleagues (2025) analyzed this systematically in large-model reasoning reinforcement learning and proposed Clip-Cov and KL-Cov to maintain entropy. Robot actions are naturally multimodal, and when RL fine-tuning a VLA, common fixes for collapse include raising the sampling temperature, widening the upper clipping bound, and adding an entropy bonus.

ExampleWhen running GRPO reinforcement learning on OpenVLA-OFT, SimpleVLA-RL strengthens exploration with three changes: dynamic sampling, raising the upper clipping bound of PPO-style clipping (following DAPO's Clip-Higher), and raising the rollout sampling temperature.

Also called
Policy Entropy Collapse, Helvetica Scenario
Related
Entropy Regularization · Exploration vs. Exploitation · Generative Adversarial Network · Action Multimodality · Group Relative Policy Optimization · SimpleVLA-RL
Sources
NIPS 2016 Tutorial: Generative Adversarial Networks (arXiv:1701.00160)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models (arXiv:2505.22617)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning (arXiv:2509.09674)

See it in the full glossary →