Embodied AI Glossary中文

Action Multimodality

动作多峰性Common

When several different actions are all correct in the same situation, so the action distribution has more than one peak.

Action multimodality means that, for the same observation, more than one action is reasonable — going around an obstacle by passing it on the left or on the right, for example. If human demonstrations contain both, the probability distribution over actions has two peaks, or modes. If a policy is trained by directly regressing a single action with mean-squared error, the model learns the average of the two peaks instead — a failure called mode averaging — which in this example might send the robot straight into the obstacle. The fix is to use a policy that can represent multimodal distributions: a Gaussian mixture model, discretizing actions into tokens and predicting them as a classification problem (as in BeT), or diffusion policies and flow matching. The Diffusion Policy paper lists handling multimodal action distributions as one of its main advantages, which is part of why so many VLA models now use a diffusion or flow-matching action head.

ExampleIn the Push-T task, demonstrators sometimes push the T-shaped block from the left and sometimes from the right. A policy trained by direct regression tends to learn the average of the two, while Diffusion Policy commits to one side and pushes through on each run.

Also called
Multimodal Action Distribution, Mode Averaging
Related
Diffusion Policy · Flow Matching · Gaussian Mixture Model · Behavior Transformer · Behavior Cloning · Diffusion Action Head
Sources
Diffusion Policy 项目页(Columbia) (Chinese)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
Behavior Transformers: Cloning k Modes with One Stone

See it in the full glossary →