Embodied AI Glossary中文

Action-free Video

无动作标签视频Common

Video that has only images, no frame-by-frame action record, like human videos and web videos.

Action-free video means an image sequence with no frame-by-frame action labels, commonly sourced from human-manipulation videos on the web, egocentric video, and robot footage where only the images were kept. It vastly outnumbers robot data with actions attached, and it carries a great deal of information about how objects move and how tasks are broken down, but it cannot be used directly for behavior cloning. Common uses fall into a few groups: training an inverse dynamics model to fill in pseudo action labels; learning discrete “latent actions” from adjacent-frame changes for pretraining, as in LAPA and Genie; first generating future frames and then inferring the action, as in UniPi and AVDC; or using it only to pretrain visual representations and world models. LAPA reports that a model pretrained this way outperformed a VLA trained with robot action labels on real-robot tasks.

ExampleAVDC trains using only action-free RGB video: it first synthesizes video of a robot completing the task with a video-generation model, then infers how the robot should move from the dense correspondence (optical flow) between adjacent frames, validated on tabletop manipulation and navigation tasks.

Also called
Actionless Video
Related
Latent Action Pretraining · Human Video Data · Internet Video Data · Pseudo Action Labels · Latent Action Model · Imitation from Observation
Sources
Ko et al. 2023: Learning to Act from Actionless Videos through Dense Correspondences (AVDC)
Ye et al. 2024: Latent Action Pretraining from Videos (LAPA)

See it in the full glossary →