EgoVLA
AdvancedA VLA pretrained on first-person human video that converts human hand motion into humanoid robot actions.
EgoVLA is a paper released in July 2025 by UC San Diego together with UIUC, MIT, NVIDIA, and others. The motivation is that real-robot data is tied to specific robot hardware and limited in scale, while first-person human video is abundant and covers a much wider range of scenes. It uses the NVILA-2B vision-language model as its backbone, first learning to predict human wrist pose and MANO hand parameters (a parametric model of the human hand) from about 500,000 human-video image-action pairs; the robot hand is then converted into that same MANO action space, and the model is fine-tuned on a small number of robot demonstrations. At deployment, wrist pose is converted into arm joint angles through inverse kinematics, and finger motion is mapped to dexterous-hand joints by a small MLP.
ExampleOn the authors' Ego Humanoid Manipulation Benchmark, built on Isaac Lab (a simulated Unitree H1 fitted with an Inspire dexterous hand, across 12 tasks), EgoVLA pretrained on human video clearly outperforms baselines trained only on robot data, with especially large gains on long-horizon and fine-grained manipulation tasks.
- Also called
- Learning Vision-Language-Action Models from Egocentric Human Videos
- Related
- Egocentric Video · Pretraining on Human Videos · MANO · Motion Retargeting · Vision-Language-Action Model · EgoScale
- Sources
- EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos (arXiv 2507.12440)
EgoVLA 项目页 (Chinese) - As of
- 2025-07