Markerless (Video-Based) Motion Capture
视频动捕(无标记动捕)AdvancedEstimating a person's 3D motion directly from ordinary video, with no markers or motion-capture suit needed.
Video-based (markerless) motion capture doesn't rely on reflective markers or inertial sensors; it uses just one or a few ordinary cameras, with a computer-vision algorithm estimating the 3D pose and motion trajectory of the human body, including the hands. Compared with optical motion capture, it needs no dedicated studio or special clothing, and can work with footage shot on a phone or pulled from the internet, so it can gather human motion data at scale; the tradeoff is lower accuracy and more noise, and monocular video adds depth and scale ambiguity plus occlusion problems, usually requiring post-processing or physical constraints to clean up. Commonly used models include WHAM and GVHMR (from Zhejiang University, which recovers human motion in world coordinates from monocular video), with HaMeR commonly used for hands. In embodied AI, this is an entry point for humanoid robots learning motion from human video and for real-time teleoperation: the estimated human motion is mapped onto the robot through motion retargeting, then used to train a motion-tracking policy.
ExampleHumanPlus uses a single RGB camera to estimate a person's body (with WHAM) and hand (with HaMeR) pose in real time, letting a humanoid robot follow the operator's movements; UC Berkeley's VideoMimic reconstructs both human motion and scene geometry from casually shot monocular video, training a humanoid robot to climb stairs, sit down, and stand up.
- Also called
- Markerless Mocap, Monocular Motion Capture, Vision-Based Motion Capture
- Related
- Motion Capture · Optical Motion Capture · Human Mesh Recovery · GVHMR · HumanPlus · VideoMimic
- Sources
- Motion capture - Wikipedia(Markerless 部分) (Chinese)
HumanPlus: Humanoid Shadowing and Imitation from Humans (arXiv 2406.10454)
VideoMimic 项目主页 (Chinese) - As of
- 2025-09