Embodied AI Glossary中文

Human Mesh Recovery

人体网格恢复HMRAdvanced

The task of estimating a person’s complete 3D body mesh — pose plus body shape — from an image or video.

Human mesh recovery estimates a person’s complete 3D body surface from an RGB image or video, rather than just a dozen or so joint points. The common approach regresses the parameters of a parametric body model like SMPL: pose parameters describe each joint’s rotation, shape parameters describe height and build, and plugging these into the model produces a human mesh. The name comes from Kanazawa, Black, Malik, and colleagues’ CVPR 2018 paper HMR, which regresses SMPL parameters directly from image pixels and uses an adversarial discriminator to constrain the result to look like a real human body. Later work, HMR 2.0 / 4D-Humans, switched to a ViT backbone and added cross-frame tracking, and methods like GVHMR further recover the person’s motion trajectory in world coordinates; HaMeR is the corresponding method for hands. In embodied AI, this is a key step for capturing full-body motion from human video and retargeting it onto a humanoid robot.

Example4D-Humans can reconstruct every person’s SMPL mesh frame by frame from a multi-person video and track each person’s identity across frames, with the result handed to a motion-retargeting tool to convert into joint trajectories for a humanoid robot.

Also called
HMR, 3D Human Pose and Shape Estimation, Human Mesh Reconstruction
Related
SMPL · GVHMR · HaMeR · Human Pose Estimation · Markerless Motion Capture · Motion Retargeting
Sources
arXiv 1712.06584: End-to-end Recovery of Human Shape and Pose(HMR, CVPR 2018)
4D-Humans / HMR 2.0 项目主页(ICCV 2023) (Chinese)

See it in the full glossary →