Variational Autoencoder
变分自编码器VAECommonA model that compresses data into a distribution over latent variables, and can sample from it to reconstruct or generate data.
The variational autoencoder was proposed by Kingma and Welling in 2013 and is a type of generative model. Its encoder compresses the input (such as an image) into a low-dimensional latent variable, but outputs a distribution (a mean and variance) rather than a single fixed point; its decoder samples from that distribution and reconstructs data from it. Training optimizes two things at once: reconstructions should look right, and the latent distribution should stay close to a standard normal distribution, which keeps the latent space continuous and smooth, so decoding a randomly sampled point still gives a plausible result. Thanks to the reparameterization trick, a model with this kind of random sampling can still be trained with gradient descent. Its most common use today is as a compressor in front of a diffusion model: Stable Diffusion and Wan both use a VAE to compress pixels into latent space before denoising; in robotics, ACT uses its conditional variant, the CVAE, to capture the variety in demonstration actions.
ExampleStable Diffusion's VAE compresses a 512×512 color image into a 64×64×4 latent variable; the diffusion model does all its denoising in that much smaller space, and only the decoder turns the result back into an image at the end.
- Also called
- VAE
- Related
- Autoencoder · Latent Space · Conditional Variational Autoencoder · Vector-Quantized Variational Autoencoder · Latent Diffusion Model · Video Tokenizer
- Sources
- Auto-Encoding Variational Bayes (arXiv 1312.6114)
High-Resolution Image Synthesis with Latent Diffusion Models (arXiv 2112.10752)