UniAct
UniAct(通用动作空间)AdvancedA method for training an embodied foundation model on a shared, discrete set of 'universal actions' common across different robots.
UniAct was released in January 2025 by Tsinghua University's Institute for AI Industry Research (AIR, Xianyuan Zhan's group) together with SenseTime, Peking University, BUPT, and the Shanghai Artificial Intelligence Laboratory, published at CVPR 2025. Different robots have very different action spaces — different numbers of joints, control schemes, and coordinate frames — so training directly on mixed cross-embodiment data causes interference. UniAct instead learns a 'universal action space': a vision-language model maps observations and instructions to a discrete universal action drawn from a vector-quantized codebook, where each code represents an atomic behavior shared across robots; a lightweight, robot-specific decoder head then translates the universal action into that robot's actual control commands. The 0.5B-parameter UniAct matches or beats the 7B OpenVLA across several real-robot and simulation evaluations; adapting to a new robot mainly just requires training a new lightweight decoder head.
ExampleAfter pretraining together on data from several robot arms such as WidowX and Franka, adapting to a new arm only requires training a small decoder head, which then translates the universal actions into that arm's control commands.
- Also called
- Universal Actions, UniAct: Universal Actions for Enhanced Embodied Foundation Models
- Related
- Unified Action Space · Cross-Embodiment · Vector Quantization · Embodiment-specific Head · OpenVLA · Latent Action
- Sources
- Universal Actions for Enhanced Embodied Foundation Models (arXiv 2501.10105)
UniAct project page - As of
- 2025-03