Embodied AI Glossary中文

VideoMimic

Advanced

A humanoid robot method that learns skills like climbing stairs and sitting down in a chair from ordinary phone-shot human videos.

VideoMimic was released by Angjoo Kanazawa, Jitendra Malik, Pieter Abbeel, and colleagues at UC Berkeley in May 2025, winning the CoRL 2025 Best Student Paper Award. Teaching a humanoid robot 'environment-aware' actions, such as climbing stairs or sitting down in a chair, requires knowing both how a human moves and what the surrounding terrain looks like. VideoMimic is a real-to-sim-to-real pipeline: from a monocular video, it simultaneously reconstructs a metrically accurate 4D human trajectory and the scene's geometry; the human motion is retargeted onto the robot, and the scene is converted into a simulator mesh; a policy that tracks these motions is then trained in simulation with reinforcement learning, and distilled into a single policy that relies only on proprioception, a height map of the terrain around the body, and a goal direction. The final policy is deployed on a 23-degree-of-freedom Unitree G1.

ExampleA phone video of a person walking up some steps and then sitting down on a bench is fed through VideoMimic, and the Unitree G1 can then autonomously climb stairs and sit down and stand up in a real environment, with the same policy working for different staircases and chairs.

Also called
Visual Imitation Enables Contextual Humanoid Control
Related
Real-to-Sim-to-Real · 4D Reconstruction · Motion Retargeting · Motion Tracking · Human Video Data · Unitree G1
Sources
Visual Imitation Enables Contextual Humanoid Control (arXiv 2505.03729)
VideoMimic project page
As of
2025-09

See it in the full glossary →