Internet Video Data
互联网视频数据CommonThe vast supply of publicly available online video — large and cheap, but without action labels a robot can use.
Internet video data refers to video from video platforms and public video datasets, mostly showing people doing chores, cooking, and using tools. Compared with robot data, it is orders of magnitude larger in scale, covers a far wider range of scenes and objects, and carries a great deal of knowledge about how objects move and how tasks break into steps. The difficulty is that it has no action labels: the frame shows a human hand rather than a robot arm, and there is no way to know what joint command each frame corresponds to. Common uses fall into three groups: pretraining visual representations (as in R3M); learning latent actions, which automatically infer an action-like code from the change between adjacent frames (as in LAPA); and training video-generation or world models that later infer actions from predicted frames. Egocentric video collections such as Ego4D are commonly used this way too.
ExampleLAPA first uses a VQ-VAE to learn discrete latent actions from adjacent frames of unlabeled video, pretrains a VLA on those latent actions, then maps them to actual robot commands using only a small amount of real-robot data.
- Also called
- Web Video Data
- Related
- Human Video Data · Action-free Video · Latent Action · LAPA · Egocentric Video · Ego4D
- Sources
- LAPA: Latent Action Pretraining from Videos (arXiv 2410.11758)
Ego4D 官网 (Chinese)