Embodied AI Glossary中文

Discrete Diffusion VLA

Discrete Diffusion VLA(离散扩散 VLA)Advanced

Decodes VLA actions with discrete diffusion: action tokens are generated in parallel, confident ones fixed first, the rest filled in over several rounds.

Discrete Diffusion VLA was released in August 2025 by researchers at the University of Hong Kong, Shanghai Jiao Tong University, and other institutions (corresponding authors Yao Mu and Ping Luo), and has been accepted at ICML 2026. Discrete VLAs like OpenVLA bin actions into tokens and generate them one at a time, autoregressively, which is slow and can't revise early mistakes; continuous diffusion-head VLAs like π0 need a separate action module bolted on. This model instead changes the action-chunk tokens in OpenVLA (built on Prismatic-7B, a Llama 2 backbone) to use bidirectional attention, and decodes them with discrete diffusion (masked prediction): all tokens start masked, each round predicts in parallel, high-confidence tokens are locked in first, the rest keep iterating, and already-filled tokens that remain uncertain can be re-masked and corrected. Action decoding shares a single Transformer and a cross-entropy objective with the language model, which also better preserves the original VLM's abilities.

ExampleOn simulation benchmarks it reaches 96.4% average success on LIBERO, 71.2% on SimplerEnv-Fractal visual matching, and 54.2% on SimplerEnv-Bridge; the authors also validated it on a real AgileX Cobot Magic platform.

Also called
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
Related
Discrete Diffusion · Vision-Language-Action Model · OpenVLA · Action Binning · Parallel Decoding · Autoregressive Decoding
Sources
Discrete Diffusion VLA (arXiv:2508.20072)
Discrete Diffusion VLA 论文 HTML 全文 v4 (Chinese)
As of
2026-05

See it in the full glossary →