Embodied AI Glossary中文

Hand Pose Estimation

手部姿态估计Common

Estimating the positions of a human hand's joints and how the fingers are bent, from an image or sensor data.

Hand pose estimation infers the positions of a human hand's joints and its finger configuration from an image, depth data, or headset sensor data. Two outputs are common: 21 2D or 3D keypoints (the wrist plus 4 points per finger), which is what MediaPipe outputs; or a parametric hand mesh, most commonly the MANO model (proposed in 2017, with 778 vertices controlled by pose and shape parameters), with HaMeR using a large vision transformer to regress a MANO mesh directly from a single image. The difficulty is that fingers are thin and prone to self-occlusion, and are further blocked by whatever object is being held. For embodied AI, this is the first step in turning human hand motion into robot motion: teleoperation gets real-time keypoints from a headset's hand tracking, which motion retargeting then converts to drive a dexterous hand; learning from human video likewise requires the hand's trajectory to be estimated first, as a pseudo-action label.

ExampleOpen-TeleVision uses an Apple Vision Pro to get the operator's hand keypoints in real time, which dex-retargeting then optimizes into dexterous-hand joint angles; when the operator makes a fist, the Unitree H1's hand follows suit.

Also called
Hand Tracking, Hand Mesh Recovery
Related
MANO · HaMeR · MediaPipe · Motion Retargeting · Hand-Object Interaction · Human Pose Estimation
Sources
HaMeR: Reconstructing Hands in 3D with Transformers (arXiv:2312.05251)
MANO 官网 (Chinese)
Open-TeleVision (arXiv:2407.01512)

See it in the full glossary →