Embodied AI Glossary中文

Discrete Diffusion

离散扩散Advanced

A diffusion model that adds and removes noise on discrete symbols, like text tokens, instead of continuous values.

An ordinary diffusion model adds Gaussian noise to continuous data — pixels, action values — but data like text or discrete action tokens can't have Gaussian noise added to it directly. Discrete diffusion instead adds noise through 'state transitions': at each step, a transition matrix randomly swaps a token for a different value, or for a special [MASK] symbol. Google's Austin and colleagues systematically formalized this framework in D3PM in 2021, and showed that using an 'absorbing state' for noising (once a token becomes MASK, it stays MASK) connects it closely to masked language models and autoregressive models. Masked diffusion has since become the dominant approach — NeurIPS 2024's MDLM, for instance, simplifies the training objective down to a mixture of masked-language-modeling losses. It's the theoretical foundation behind diffusion language models and discrete-diffusion-style VLAs.

ExampleDiscrete Diffusion VLA discretizes an action chunk into tokens, masks all of them, then progressively reveals the easiest ones first based on confidence, re-masking and recomputing the uncertain positions, and reports a 96.4% average success rate on LIBERO.

Also called
Masked Diffusion
Related
Diffusion Model · Diffusion Language Model · Parallel Decoding · Discrete Diffusion VLA · Denoising Diffusion Probabilistic Model · Action Binning
Sources
Structured Denoising Diffusion Models in Discrete State-Spaces (D3PM, arXiv:2107.03006)
Simple and Effective Masked Diffusion Language Models (MDLM, arXiv:2406.07524)
Discrete Diffusion VLA (arXiv:2508.20072)

See it in the full glossary →