Latent Action Pretraining
潜在动作预训练AdvancedExtracting unlabeled “latent actions” from video, then pretraining a robot model to predict them.
Latent action pretraining first learns “latent actions” — an encoding of what changed between two adjacent frames — from video that has no action labels, then pretrains a VLA to predict these latent actions; a small amount of real-robot data is used at the end to map the latent actions onto real robot actions. The representative work is LAPA (October 2024), from researchers at the University of Washington, KAIST, Microsoft Research, NVIDIA, and others, which learns discrete latent actions with a VQ-VAE (a vector-quantized autoencoder). Its value is that it can exploit huge amounts of internet and human-manipulation video that carries no robot action labels at all. Genie's latent action model, UniVLA, and AgiBot's GO-1 all take a similar approach.
ExampleThe LAPA paper reports positive transfer even when pretraining only on Something-Something V2 human-manipulation videos; on real-robot tasks that require language conditioning and generalization, it beats OpenVLA, which was trained with real action labels, while using roughly one-thirtieth the pretraining compute.
- Related
- Latent Action · Latent Action Model · LAPA · Action-free Video · Pretraining on Human Videos · Vector-Quantized Variational Autoencoder
- Sources
- Latent Action Pretraining from Videos (LAPA, arXiv:2410.11758)
LAPA 论文 HTML 版 (Chinese) - As of
- 2024-10