Embodied AI Glossary中文

Conditional Variational Autoencoder

条件变分自编码器CVAECommon

A VAE that learns an output distribution given a condition, so the same input can generate several valid outputs.

A conditional variational autoencoder extends the variational autoencoder (VAE): both the encoder and decoder additionally take a condition, such as the current observation, and the model learns the distribution of outputs given that condition, rather than one single answer. It was proposed by Sohn, Lee, and Yan at NeurIPS 2015, originally for structured-output prediction tasks such as image segmentation. During training, the encoder sees the real output, compresses its “style” into a latent variable z, and the decoder reconstructs the output from the condition and z; at inference, z is instead drawn from the prior. Imitation learning uses it to handle the variety in demonstrations: for the same scene, different people might take different paths, and ordinary regression would average several valid approaches into one wrong action. ACT (2023) has the encoder take the joint positions and target action sequence to get z, and at test time drops the encoder and just sets z to the prior mean of 0.

ExampleACT uses a CVAE to predict a 100-step action chunk at once, and with only about 10 minutes of demonstration data, performs fine-motor tasks on a low-cost dual-arm setup with an 80% to 90% success rate.

Also called
CVAE, Conditional VAE
Related
Variational Autoencoder · Action Chunking with Transformers · Action Multimodality · Action Chunking · Generative Model · Autoencoder
Sources
Learning Structured Output Representation using Deep Conditional Generative Models (NeurIPS 2015)
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv:2304.13705)

See it in the full glossary →