Embodied AI Glossary中文

Motus

Advanced

A latent-action world model that combines understanding, video generation, and action into three experts inside one model.

Motus was released in December 2025 by Jun Zhu's group at Tsinghua University together with Peking University and Horizon Robotics, with code and weights open-sourced under Apache 2.0. Earlier approaches often split understanding, a world model (predicting future frames), and control into separate models, making it hard to jointly exploit large-scale heterogeneous data. Motus uses a Mixture-of-Transformers (MoT) architecture to connect an understanding expert, a video-generation expert, and an action expert together — the video part is built on Wan2.2-5B and the understanding part on Qwen3-VL-2B, about 8 billion parameters in total — and uses UniDiffuser-like scheduling to switch among modes such as world model, VLA, inverse dynamics model, and video generation. It learns latent actions from optical flow, letting video with no action labels also join pretraining, paired with three-stage training and a six-tier data pyramid. Across 50 tasks on RoboTwin 2.0, it reaches an average success rate of about 87%, which the paper reports as 15% higher than X-VLA and 45% higher than π0.5.

ExampleThe same Motus model: given the current frame and an instruction, it outputs an action directly (VLA mode); given a frame and an action, it predicts the following video (world-model mode); given a before-and-after pair of frames, it infers what action happened in between (inverse-dynamics mode).

Also called
A Unified Latent Action World Model
Related
World Action Model · Latent Action · Mixture-of-Transformers · Inverse Dynamics Model · Data Pyramid · RoboTwin
Sources
Motus: A Unified Latent Action World Model (arXiv 2512.13030)
thu-ml/Motus (GitHub)
Motus project page
As of
2025-12

See it in the full glossary →