PH2D
PH2D 数据集(HAT)AdvancedFirst-person human manipulation data collected with a VR headset, co-trained with humanoid robot data to train one policy.
PH2D (Physical Human-Humanoid Data) comes from the paper Humanoid Policy ~ Human Policy, led by UC San Diego together with CMU, MIT, Apple, and others, published at CoRL 2025. The authors had people perform manipulation tasks directly while wearing consumer VR headsets like Apple Vision Pro or Meta Quest 3, automatically recording first-person video plus the 3D position of the head, wrists, and fingertips — about 27,000 language-annotated demonstrations in total. The companion model, HAT (Human Action Transformer), places humans and humanoid robots into one shared state-action space, and its output is then retargeted onto the robot's joints. The underlying idea: a person collecting their own data is far faster than teleoperating a robot, and as long as the representation is aligned, this data can be co-trained with a smaller amount of robot data to improve generalization and robustness.
ExampleA person wearing a Vision Pro quickly records a large number of pick-and-place demonstrations, which are combined with a smaller set of Unitree H1 teleoperation data to train HAT; the resulting policy is deployed directly on the H1.
- Also called
- HAT (Human Action Transformer), Humanoid Policy ~ Human Policy
- Related
- Egocentric Video · Human Video Data · Co-training · Embodiment Gap · Open-TeleVision · Unitree H1
- Sources
- Humanoid Policy ~ Human Policy (arXiv 2503.13441)
Humanoid Policy ~ Human Policy 项目主页 (Chinese) - As of
- 2025-09