Tencent HY-Embodied (Hunyuan Embodied)
腾讯 HY-Embodied(混元具身)HY-EmbodiedAdvancedAn open-source embodied foundation model series from Tencent Robotics X and the Hunyuan team, including both an embodied VLM and a VLA.
This is an open-source embodied foundation model family from Tencent's Robotics X Lab and its Hunyuan vision team. HY-Embodied-0.5, from April 2026, is a vision-language model built for robots, strengthening spatial and temporal perception and embodied reasoning, using a Mixture-of-Transformers (MoT) architecture, with a 2B-active (4B total) on-device version and a 32B version — the MoT-2B was open-sourced. The same month's 0.5-X continues post-training on top of it, focused on task planning and risk judgment. June's Hy-Embodied-0.5-VLA adds a flow-matching action expert on this backbone, trained on more than 10,000 hours of bimanual data collected with the company's own fingertip UMI device, with over 2,000 hours of that open-sourced; in July, a MoE-architecture VLM-1.0 followed (about 3B active / 30B total parameters). The project page sits on Tencent's Tairos embodied-AI open platform.
ExampleHy-Embodied-0.5-VLA reports success rates of 90.9% (Clean) and 90.1% (Randomized) on the RoboTwin 2.0 simulation benchmark.
- Also called
- HY-Embodied-0.5, HY-Embodied-0.5-X, Hy-Embodied-0.5-VLA, HY-VLA-0.5, Hy-Embodied-VLM-1.0
- Related
- Embodied Foundation Model · Vision-Language Model · Vision-Language-Action Model · Mixture-of-Transformers · Universal Manipulation Interface · Tencent Tairos Embodied AI Open Platform
- Sources
- Tencent-Hunyuan/HY-Embodied (GitHub)
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents (arXiv 2604.07430)
Tencent-Hunyuan/Hy-Embodied-0.5-VLA (GitHub) - As of
- 2026-07