OmniH2O
CommonCMU's 2024 humanoid whole-body teleoperation system: VR, cameras, voice, or GPT-4o can all drive the same robot.
OmniH2O is a humanoid whole-body teleoperation and learning system from Tairan He, Zhengyi Luo, Guanya Shi, and colleagues at Carnegie Mellon University and Shanghai Jiao Tong University, proposed in June 2024 and published at CoRL 2024 as an upgrade to the same group's H2O. Its core idea is to use “kinematic pose” — the position and orientation of key body parts — as a unified control interface: whether a command comes from a VR headset, an RGB camera, voice, or GPT-4o, it's first converted into a target pose, which the same whole-body control policy then tracks. That policy is trained in simulation with reinforcement learning: after large-scale retargeting and augmentation of AMASS human motion data, a teacher policy is trained using privileged information (full state only available in simulation), then distilled into a student policy that uses only sparse sensor inputs and can run on the real robot. The platform is a Unitree H1 fitted with dexterous hands. The team also released OmniH2O-6, the first whole-body-control dataset for humanoids covering 6 everyday tasks, and used it to train autonomous skills with Diffusion Policy.
ExampleAutonomous shadow-boxing: the robot's head-camera feed is sent to GPT-4o with a prompt saying to throw a left jab at a blue target, a right jab at a red target, and stay still if there's no target; GPT-4o answers only A, B, or C each time, and OmniH2O's whole-body policy carries out the chosen action.
- Also called
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning, Omni Human-to-Humanoid
- Related
- H2O · Whole-Body Teleoperation · Teacher-Student Distillation · Privileged Information · Motion Retargeting · HumanPlus
- Sources
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning (arXiv 2406.08858)
OmniH2O 项目主页 (Chinese) - As of
- 2024-11