Latent Diffusion Model
潜在扩散模型LDMAdvancedA model that first compresses data into a low-dimensional latent space with an autoencoder, then runs diffusion generation there.
A diffusion model generates data by denoising step by step, but doing this directly on pixels is computationally expensive for high-resolution images and video. In December 2021, Rombach and colleagues at Germany's CompVis group proposed the latent diffusion model (CVPR 2022): an autoencoder is trained first to compress an image into a much smaller latent variable, the diffusion model denoises only in that latent space, and a decoder reconstructs the image at the end; conditions such as text are injected via cross-attention. This dramatically cuts training and inference cost, and the open-source Stable Diffusion is built on this framework. Video generation and world models commonly follow the same idea — for example, the diffusion-based world foundation models in NVIDIA's Cosmos run inside the latent space of the Cosmos video tokenizer (an encoder that compresses video into a latent variable), at a spatial-temporal compression ratio of 8×8×8.
ExampleStable Diffusion first compresses a 512×512 color image into a 64×64×4 latent variable, denoises on this much smaller tensor, and only decodes back to a 512×512 image at the end.
- Also called
- LDM, Latent Diffusion
- Related
- Diffusion Model · Variational Autoencoder · Diffusion Transformer · Video Tokenizer · NVIDIA Cosmos · Cross-Attention
- Sources
- High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752)
Cosmos World Foundation Model Platform for Physical AI (arXiv:2501.03575) - As of
- 2025-01