Latent Action Model
潜在动作模型LAMCommonThe network that produces latent actions from unlabeled video, trained by having an encoder-decoder pair reconstruct the next frame.
A latent action model is the network that produces latent actions. The typical design comes from DeepMind's 2024 Genie: an encoder looks at the current frame and the next frame and outputs a latent action, playing the role of an inverse dynamics model; a decoder then takes only the current frame plus that latent action and reconstructs the next frame, playing the role of a forward dynamics model. In between, vector quantization (the VQ-VAE approach) restricts the latent action to a very small discrete codebook — Genie uses just 8 codes — so the latent action can't encode a whole image, only ‘what changed.’ Once trained, the encoder can attach latent-action labels to huge amounts of video. A common difficulty is that task-irrelevant changes, such as camera shake or a person walking through the background, also get encoded; UniVLA addresses this by learning in DINO feature space instead and using the language instruction to strip out such distractions.
ExampleLAPA has three stages: first train a latent-action quantization model on video, then train a VLA to predict latent actions from images and instructions, and finally fine-tune on a small amount of robot data to replace latent actions with real ones.
- Also called
- LAM
- Related
- Latent Action · Inverse Dynamics Model · Forward Dynamics Model · Vector-Quantized Variational Autoencoder · Genie (Original) · UniVLA
- Sources
- Genie: Generative Interactive Environments (arXiv 2402.15391)
Latent Action Pretraining from Videos (LAPA, arXiv 2410.11758)
UniVLA: Learning to Act Anywhere with Task-centric Latent Actions (arXiv 2505.06111)