Action-free Video
无动作标签视频CommonVideo that has only images, no frame-by-frame action record, like human videos and web videos.
Action-free video means an image sequence with no frame-by-frame action labels, commonly sourced from human-manipulation videos on the web, egocentric video, and robot footage where only the images were kept. It vastly outnumbers robot data with actions attached, and it carries a great deal of information about how objects move and how tasks are broken down, but it cannot be used directly for behavior cloning. Common uses fall into a few groups: training an inverse dynamics model to fill in pseudo action labels; learning discrete “latent actions” from adjacent-frame changes for pretraining, as in LAPA and Genie; first generating future frames and then inferring the action, as in UniPi and AVDC; or using it only to pretrain visual representations and world models. LAPA reports that a model pretrained this way outperformed a VLA trained with robot action labels on real-robot tasks.
ExampleAVDC trains using only action-free RGB video: it first synthesizes video of a robot completing the task with a video-generation model, then infers how the robot should move from the dense correspondence (optical flow) between adjacent frames, validated on tabletop manipulation and navigation tasks.
- Also called
- Actionless Video
- Related
- Latent Action Pretraining · Human Video Data · Internet Video Data · Pseudo Action Labels · Latent Action Model · Imitation from Observation
- Sources
- Ko et al. 2023: Learning to Act from Actionless Videos through Dense Correspondences (AVDC)
Ye et al. 2024: Latent Action Pretraining from Videos (LAPA)