LAPA
CommonPretrains a VLA on action-label-free video by first learning latent actions, then mapping them to real actions with little robot data.
LAPA was proposed in October 2024 by researchers from KAIST, the University of Washington, Microsoft Research, NVIDIA, and the Allen Institute for AI, published at ICLR 2025. Pretraining a VLA normally requires action labels collected through teleoperation, which limits both data sources and scale; LAPA instead aims to use the huge amount of action-label-free video already on the web. It works in three steps. First, a VQ-VAE-style objective trains a quantization model that encodes the change between two consecutive frames into a discrete latent action. Second, a vision-language model (a 7B LWM) is pretrained to predict these latent actions from images and a task description. Finally, it's fine-tuned on a small amount of real-robot data that maps the latent actions to real robot actions. The paper reports that on real-robot tasks requiring language understanding and generalization, it beats OpenVLA — which was trained with real action labels — by 6.22%, while being more than 30x more pretraining-efficient.
ExamplePretraining latent actions on only about 220,000 everyday human manipulation videos from Something-Something V2, then fine-tuning on a small number of robot trajectories, already outperforms OpenVLA pretrained on the Bridge robot dataset.
- Also called
- LAPA: Latent Action Pretraining from Videos, Latent Action Pretraining for General Action Models
- Related
- Latent Action · Latent Action Pretraining · Latent Action Model · Vector-Quantized Variational Autoencoder · Action-free Video · OpenVLA
- Sources
- Latent Action Pretraining from Videos (arXiv 2410.11758)
LAPA 项目主页 (Chinese) - As of
- 2025-04