Entropy Collapse / Mode Collapse
熵坍缩 / 模式坍缩AdvancedA model's output diversity collapsing, so it only ever produces a handful of answers or actions.
These two terms describe the same broad phenomenon. Mode collapse was first used for GANs (generative adversarial networks): Goodfellow's 2016 tutorial describes it as the generator mapping many different input noise vectors to the same output, covering only one or two “modes” of the data distribution. Entropy collapse is used more in reinforcement learning: a policy's entropy (how random it is) drops rapidly early in training, the model settles on almost one single behavior, stops exploring, and performance plateaus; Cui and colleagues (2025) analyzed this systematically in large-model reasoning reinforcement learning and proposed Clip-Cov and KL-Cov to maintain entropy. Robot actions are naturally multimodal, and when RL fine-tuning a VLA, common fixes for collapse include raising the sampling temperature, widening the upper clipping bound, and adding an entropy bonus.
ExampleWhen running GRPO reinforcement learning on OpenVLA-OFT, SimpleVLA-RL strengthens exploration with three changes: dynamic sampling, raising the upper clipping bound of PPO-style clipping (following DAPO's Clip-Higher), and raising the rollout sampling temperature.
- Also called
- Policy Entropy Collapse, Helvetica Scenario
- Related
- Entropy Regularization · Exploration vs. Exploitation · Generative Adversarial Network · Action Multimodality · Group Relative Policy Optimization · SimpleVLA-RL
- Sources
- NIPS 2016 Tutorial: Generative Adversarial Networks (arXiv:1701.00160)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models (arXiv:2505.22617)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning (arXiv:2509.09674)