HumanPlus
CommonStanford's 2024 humanoid system: one RGB camera lets a robot shadow a person in real time, then learn skills from it.
HumanPlus is a humanoid data-collection and learning system proposed in June 2024 by Zipeng Fu, Qingqing Zhao, Qi Wu, Gordon Wetzstein, and Chelsea Finn at Stanford, published at CoRL 2024 and a finalist for best paper. Its hardware is a Unitree H1 fitted with two 6-DOF Inspire dexterous hands, for 33 degrees of freedom total. It works in two steps. First, a low-level controller called the Humanoid Shadowing Transformer is trained in simulation with reinforcement learning on about 40 hours of human motion data from AMASS; once deployed, a single RGB camera is enough to estimate the operator's body and hand pose, letting the robot “shadow” them in real time. This teleoperation method is then used to collect demonstrations, training a Humanoid Imitation Transformer, built on ACT, that acts autonomously using first-person video from a stereo RGB camera on the robot's head. The work shows that whole-body teleoperation doesn't necessarily require expensive motion-capture equipment.
ExampleStanding up and walking after putting on shoes: using at most 40 shadowing-collected demonstrations per task, the trained robot completed it autonomously with 60% success; unloading a warehouse shelf reached 90%, and typing reached 80%.
- Also called
- HumanPlus: Humanoid Shadowing and Imitation from Humans
- Related
- Humanoid Robot · Whole-Body Teleoperation · Unitree H1 · Action Chunking with Transformers · AMASS (Archive of Motion Capture as Surface Shapes) · OmniH2O
- Sources
- HumanPlus: Humanoid Shadowing and Imitation from Humans (arXiv 2406.10454)
HumanPlus 项目主页 (Chinese) - As of
- 2024-11