Embodied AI Glossary中文

Pretraining on Human Videos

人类视频预训练Common

Pretraining a robot model on large amounts of video of humans doing things, then fine-tuning it with a small amount of robot data.

Real-robot data is expensive to collect and stays small in scale, while video of humans going about everyday tasks — first-person datasets like Ego4D, or general internet video — is abundant and covers far more scenes, so a common approach is pretraining on human video first and fine-tuning with a small amount of robot data. Two difficulties stand out: videos have no robot action labels, and human hands and robot grippers have different structures (an embodiment gap). There are three main approaches. One learns only a visual representation, as R3M (CoRL 2022) does, training a visual encoder on Ego4D with time-contrastive learning and video-language alignment. A second extracts a latent action as a pseudo-label from the change between consecutive frames, as LAPA does, using a VQ-VAE to get discrete latent actions for pretraining a VLA. A third estimates the human hand's 3D motion directly as the action, as Being-H0 does, treating the human hand as a general-purpose manipulator to pretrain a VLA.

ExampleR3M pretrains a visual representation on human Ego4D videos, freezes it, and attaches a small policy network; with the representation frozen, a Franka arm learns a variety of manipulation tasks in a real, cluttered apartment from just 20 demonstrations.

Also called
Human Video Pretraining, Learning from Human Videos
Related
Human Video Data · Egocentric Video · Latent Action Pretraining · Embodiment Gap · R3M · LAPA
Sources
R3M: A Universal Visual Representation for Robot Manipulation
Latent Action Pretraining from Videos (LAPA)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
As of
2025-07

See it in the full glossary →