Embodied AI Glossary中文

X-VLA

Advanced

A 0.9B cross-embodiment VLA that unifies many robots by giving each one its own learnable 'soft prompt.'

X-VLA was proposed in October 2025 by Tsinghua University's Institute for AI Industry Research (AIR) together with the Shanghai Artificial Intelligence Laboratory and others, and has been accepted at ICLR 2026. The difficulty in cross-embodiment training is that different robots have very different cameras and action spaces, so training on them mixed together causes interference. X-VLA instead gives each data source its own set of learnable embedding vectors as a 'soft prompt' that tells the model which hardware it's currently dealing with, while sharing the rest of the backbone entirely; actions are generated with flow matching, and the backbone is a standard Transformer encoder. The 0.9-billion-parameter version is pretrained on 290,000 trajectories combined from DROID, RoboMIND, and AgiBot, and leads on 6 simulation benchmarks — including LIBERO, SimplerEnv, and CALVIN — and 3 real-robot platforms. The code is open-sourced under Apache-2.0 and has been integrated into LeRobot.

ExampleOn a real-robot clothes-folding task (Soft-Fold), X-VLA-0.9B is reported to reach 100% success, folding 33 garments per hour.

Also called
X-VLA-0.9B, X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
Related
Cross-Embodiment · Prompt Tuning / Soft Prompt · Flow Matching · Vision-Language-Action Model · LeRobot · Cross-Embodiment Data
Sources
X-VLA (arXiv:2510.10274)
2toinf/X-VLA (GitHub)
As of
2026-09

See it in the full glossary →