Embodied AI Glossary中文

XR-1

北京人形 XR-1Advanced

A VLA foundation model from X-Humanoid that pretrains across robot embodiments using a 'unified vision-motion encoding.'

XR-1 was released by the Beijing Humanoid Robot Innovation Center (X-Humanoid) in November 2025 (some authors also hold positions at Peking University and Beihang University), with the paper accepted as an oral presentation at ICML 2026. Its core idea is UVMC (Unified Vision-Motion Coding): a dual-branch vector-quantized variational autoencoder encodes 'how the image changes' and 'how the robot moves' into the same discrete codebook, with an alignment loss keeping the two consistent. Training has three stages: self-supervised learning of UVMC; pretraining on about 164 million frames of cross-embodiment data (Open-X, the team's own XR-D, RoboMIND, and Ego4D human video); and finally task-specific fine-tuning. The main model reuses the π0 architecture, with a lighter version, XR-1-Light, based on Florence-2. Code and weights are open-sourced on GitHub, Hugging Face, and ModelScope.

ExampleThe authors ran more than 14,000 real-robot trials across six embodiments — Tiangong 1.0/2.0, single- and dual-arm UR5e, dual-arm Franka, and AgileX Cobot Magic — covering more than 120 manipulation tasks.

Also called
X-Humanoid XR-1, XR-1-Light
Related
Cross-Embodiment · Vector-Quantized Variational Autoencoder · Latent Action · Beijing Humanoid Robot Innovation Center · Tiangong · π0
Sources
XR-1 (arXiv:2511.02776)
XR-1 论文 HTML 版 (Chinese)
As of
2026-09

See it in the full glossary →