Embodied AI Glossary中文

HybridVLA

Advanced

A VLA that predicts actions both by diffusion and autoregressively, inside the same large language model.

HybridVLA was released in March 2025 by Shanghang Zhang's group at Peking University, together with the Beijing Academy of Artificial Intelligence, CUHK, and Fudan University. Autoregressive VLA models discretize actions into tokens, which lets them inherit a vision-language model's reasoning ability, but discretization breaks the continuity of actions and hurts fine motor control; diffusion VLA models instead attach a separate diffusion head to output continuous actions, but only use the features a VLM extracts, without benefiting from token-by-token generative reasoning. HybridVLA embeds diffusion denoising directly into a large language model's next-token-prediction process, so one model produces both a diffusion action and an autoregressive action at once, which are then adaptively fused through “collaborative action ensembling.” The model is built on Prismatic 7B (a LLaMA-2 backbone); the paper reports average success rates 14% and 19% higher than the previous best methods on simulation and real-robot tasks, respectively.

ExampleAt every step, the same HybridVLA model produces two action predictions, one diffusion-based and one autoregressive, which are fused before being handed to the robot arm to execute.

Also called
Hybrid-VLA, Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Related
Vision-Language-Action Model · Hybrid Autoregressive-Diffusion Architecture · Diffusion Action Head · Action Binning · CogACT · OpenVLA
Sources
HybridVLA (arXiv 2503.10631)
HybridVLA project page
PKU-HMI-Lab/Hybrid-VLA (GitHub)
As of
2025-06

See it in the full glossary →