Latent Space
潜在空间CommonThe internal representation space a model compresses raw data into, where similar things end up close together.
Latent space is the vector space a neural network uses internally to represent data. An encoder compresses high-dimensional raw data — images, audio, actions — into vectors with far fewer dimensions, and the space those vectors live in is the latent space; when training goes well, samples with similar meaning end up close together in it. Its value is that it saves compute and captures what matters: generating, predicting, or planning in latent space is much cheaper than working directly with pixels, and it's less easily thrown off by irrelevant detail. Autoencoders and variational autoencoders (VAEs, autoencoders that can sample to generate new data) both rely on it. Common uses in embodied AI include latent diffusion models, which denoise and generate images and video in latent space; latent world models, which predict the future in terms of latent states; and latent action models, which learn an abstract ‘action code’ from video with no action labels.
ExampleA latent diffusion model (the basis of Stable Diffusion) first uses a pretrained autoencoder to compress an image into a much smaller latent representation, runs denoising diffusion in that space, and only decodes back to pixels at the end — far cheaper to train and run than diffusing directly on pixels.
- Also called
- Latent Feature Space
- Related
- Embedding · Variational Autoencoder · Latent Diffusion Model · Latent World Model · Latent Action · Representation Learning
- Sources
- Wikipedia: Latent space
High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752)