Embodied AI Glossary中文

Latent Diffusion Model

潜在扩散模型LDMAdvanced

A model that first compresses data into a low-dimensional latent space with an autoencoder, then runs diffusion generation there.

A diffusion model generates data by denoising step by step, but doing this directly on pixels is computationally expensive for high-resolution images and video. In December 2021, Rombach and colleagues at Germany's CompVis group proposed the latent diffusion model (CVPR 2022): an autoencoder is trained first to compress an image into a much smaller latent variable, the diffusion model denoises only in that latent space, and a decoder reconstructs the image at the end; conditions such as text are injected via cross-attention. This dramatically cuts training and inference cost, and the open-source Stable Diffusion is built on this framework. Video generation and world models commonly follow the same idea — for example, the diffusion-based world foundation models in NVIDIA's Cosmos run inside the latent space of the Cosmos video tokenizer (an encoder that compresses video into a latent variable), at a spatial-temporal compression ratio of 8×8×8.

ExampleStable Diffusion first compresses a 512×512 color image into a 64×64×4 latent variable, denoises on this much smaller tensor, and only decodes back to a 512×512 image at the end.

Also called
LDM, Latent Diffusion
Related
Diffusion Model · Variational Autoencoder · Diffusion Transformer · Video Tokenizer · NVIDIA Cosmos · Cross-Attention
Sources
High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752)
Cosmos World Foundation Model Platform for Physical AI (arXiv:2501.03575)
As of
2025-01

See it in the full glossary →