Diffusion Action Head
扩散动作头CommonAn output module attached after a policy's backbone that generates continuous actions through diffusion denoising.
A diffusion action head is a type of action head, the part of a policy network that produces the final action. A backbone, such as a Transformer or VLM, first encodes images, language, and robot state into features; the diffusion head then conditions on those features and, starting from random noise, denoises over multiple steps to generate a segment of continuous action. Compared with direct regression, it can express a multimodal action distribution: when a demonstration set shows both “go around the left” and “go around the right” as valid, it doesn't average them into a middle path that runs into the obstacle; compared with discretizing actions into tokens, it keeps the precision of continuous values. The cost is that inference needs multiple network calls. It scales from small to large: Octo uses a 3-layer MLP, while GR00T N1 and CogACT use a dedicated diffusion Transformer as their action module; π0's flow-matching action expert works on a similar principle and is often discussed alongside it.
ExampleOcto's Transformer backbone outputs the embedding of a readout token, which is handed to a 3-layer-MLP diffusion head that denoises over 20 steps to generate a segment of action. The paper's comparison found the robot moved hesitantly with an MSE regression head, and imprecisely, often grasping at empty air, with a discrete action head.
- Also called
- Diffusion Head
- Related
- Action Head · Action Expert · Diffusion Policy · Diffusion Transformer · Action Multimodality · Continuous Action Regression
- Sources
- Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action (arXiv 2411.19650)