Diffusion Language Model
扩散语言模型dLLMAdvancedA language model that denoises a whole span of text in parallel using diffusion, instead of generating it word by word.
A diffusion language model applies the diffusion idea to text: generation starts from a span of tokens that are all masked or noised, and predicts them together over several steps, gradually filling in content — each step can fill in several tokens at once, and can also revise earlier guesses. Representative work includes LLaDA (2025, from a Renmin University-led team), whose 8B version matches LLaMA3 8B on in-context learning; Inception Labs' commercial model Mercury; and Google DeepMind's experimental Gemini Diffusion, which the company reports sampling at about 1,479 tokens per second. Its appeal is fast parallel decoding and the ability to use context in both directions at once. Embodied AI has picked up the same idea for decoding actions in parallel too, as in Discrete Diffusion VLA.
ExampleWhen LLaDA answers a question, it first generates a whole span of masked tokens, predicts all the masked positions at each step, keeps the ones it's confident about, re-masks the uncertain ones, and repeats over several rounds until everything is filled in.
- Also called
- dLLM, Diffusion LLM, Masked Diffusion Language Model
- Related
- Discrete Diffusion · Large Language Model · Parallel Decoding · Autoregressive Decoding · Diffusion Model · Discrete Diffusion VLA
- Sources
- Large Language Diffusion Models (LLaDA, arXiv:2502.09992)
Gemini Diffusion - Google DeepMind
Mercury: Ultra-Fast Language Models Based on Diffusion (arXiv:2506.17298) - As of
- 2026-09