Pretraining on Human Videos
人类视频预训练CommonPretraining a robot model on large amounts of video of humans doing things, then fine-tuning it with a small amount of robot data.
Real-robot data is expensive to collect and stays small in scale, while video of humans going about everyday tasks — first-person datasets like Ego4D, or general internet video — is abundant and covers far more scenes, so a common approach is pretraining on human video first and fine-tuning with a small amount of robot data. Two difficulties stand out: videos have no robot action labels, and human hands and robot grippers have different structures (an embodiment gap). There are three main approaches. One learns only a visual representation, as R3M (CoRL 2022) does, training a visual encoder on Ego4D with time-contrastive learning and video-language alignment. A second extracts a latent action as a pseudo-label from the change between consecutive frames, as LAPA does, using a VQ-VAE to get discrete latent actions for pretraining a VLA. A third estimates the human hand's 3D motion directly as the action, as Being-H0 does, treating the human hand as a general-purpose manipulator to pretrain a VLA.
ExampleR3M pretrains a visual representation on human Ego4D videos, freezes it, and attaches a small policy network; with the representation frozen, a Franka arm learns a variety of manipulation tasks in a real, cluttered apartment from just 20 demonstrations.
- Also called
- Human Video Pretraining, Learning from Human Videos
- Related
- Human Video Data · Egocentric Video · Latent Action Pretraining · Embodiment Gap · R3M · LAPA
- Sources
- R3M: A Universal Visual Representation for Robot Manipulation
Latent Action Pretraining from Videos (LAPA)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos - As of
- 2025-07