Embodied AI Glossary中文

Vision-Language-Latent-Action

ViLLA 架构ViLLAAdvanced

The hierarchical architecture behind AgiBot's GO-1: it first predicts latent action tokens, then decodes them into real robot actions.

ViLLA is the framework AgiBot introduced in March 2025 alongside GO-1 (Genie Operator-1), its general-purpose embodied foundation model. A standard VLA (vision-language-action model) maps images and instructions directly to actions; ViLLA inserts an intermediate “latent action” layer with three parts. A latent action model (LAM) trains on human videos (such as Ego4D) and robot trajectories, using inverse and forward dynamics to compress the change between two adjacent frames into discrete latent action tokens. A latent planner, built on an InternVL2.5-2B backbone, predicts these tokens from multi-view images and the instruction. An action expert then decodes them into continuous, low-level action chunks using a diffusion objective. The benefit is that human videos with no action labels can still be used for pretraining, improving data efficiency. GO-1 was open-sourced in September 2025, together with a lighter GO-1 Air variant that drops the latent planner.

ExampleWhen GO-1 carries out a tabletop instruction, the latent planner first looks at three camera views plus the instruction and outputs a handful of latent action tokens — an abstract sketch of “what to do next.” The action expert then denoises conditioned on those tokens to generate the next 30 timesteps of continuous action.

Also called
ViLLA, Latent Planner + Action Expert
Related
AgiBot GO-1 · Latent Action Model · Latent Action · Action Expert · Hierarchical Architecture · Latent Action Pretraining
Sources
AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems (arXiv:2503.06669)
OpenDriveLab/AgiBot-World (GitHub, GO-1 / GO-1 Air 开源) (Chinese)
As of
2025-09

See it in the full glossary →