Embodied AI Glossary中文

Autoencoder

自编码器AECommon

A network that compresses input into a short vector and then reconstructs it, learning what matters most in the data.

An autoencoder consists of an encoder and a decoder: the encoder compresses input, such as an image, into a low-dimensional vector, and the decoder reconstructs the input from that vector, trained so the reconstruction is as close as possible to the original, with no human labels needed. Because the middle vector is smaller than the input, the network is forced to keep only the most important information, which is why it's commonly used for dimensionality reduction and feature learning. The idea goes back decades in neural networks, to work such as LeCun's in 1987. Common variants include the variational autoencoder (VAE), which makes the middle vector follow a probability distribution so new samples can be generated; the masked autoencoder (MAE), which hides part of the input and reconstructs it, used for visual pretraining; and VQ-VAE, which replaces the vector with a discrete index into a codebook. In embodied AI, latent diffusion models, video tokenizers, and latent action models all use this to compress high-dimensional data into a latent space before processing it.

Example“World Models” used a convolutional VAE to compress every frame of a racing-game screen into a 32-dimensional vector, so the prediction model downstream only ever operates on those 32 numbers.

Also called
AE
Related
Variational Autoencoder · Conditional Variational Autoencoder · Masked Autoencoder · Vector-Quantized Variational Autoencoder · Latent Space · Latent Diffusion Model
Sources
Deep Learning, Chapter 14: Autoencoders (Goodfellow, Bengio, Courville)
World Models (Ha & Schmidhuber, interactive article)

See it in the full glossary →