Embodied AI Glossary中文

Being-H0

智在无界 Being-H0Advanced

BeingBeyond's series of dexterous-manipulation VLA models, pretrained on large-scale video of human hands doing manipulation.

Being-H0 is a vision-language-action model released in July 2025 by Peking University, Renmin University of China, and the embodied-AI company BeingBeyond (智在无界), later accepted at ICML 2026. Its idea is to treat the human hand as a universal manipulator: it first pretrains on human hand-manipulation data, using a part-specific action tokenizer to compress wrist and finger motion into tokens, then does 3D spatial alignment, and finally post-trains on a small amount of robot data to transfer onto a dexterous robot hand. Its companion dataset, UniHand, pools motion capture, VR, and ordinary video, totaling about 1,100 hours and 165 million instruction samples. Being-H0.5, from January 2026, expands the data to more than 35,000 hours across 30 embodiments, with an emphasis on cross-embodiment transfer; Being-H0.7, from April, switches to a latent-space world-action model.

ExampleIn real-robot experiments, the researchers post-trained Being-H0 onto a Franka arm fitted with an Inspire six-degree-of-freedom dexterous hand for dexterous manipulation tasks.

Also called
Being-H0.5, Being-H0.7, Being-H Series
Related
Pretraining on Human Videos · UniHand · Vision-Language-Action Model · Dexterous Manipulation · Cross-Embodiment · BeingBeyond
Sources
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos (arXiv 2507.15597)
Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization (arXiv 2601.12993)
BeingBeyond/Being-H0 GitHub
As of
2026-05

See it in the full glossary →