WMPO
WMPO(基于世界模型的策略优化)AdvancedA method that lets a VLA do reinforcement learning on trajectories 'imagined' by a video world model, with no real-robot interaction.
WMPO was proposed in November 2025 by researchers at HKUST and ByteDance Seed. A VLA trained purely on expert demonstrations never learns to correct itself after a failure, while doing reinforcement learning directly on a real robot is too sample-expensive. WMPO first trains a video world model that predicts pixel-level frames, lets the VLA repeatedly practice inside trajectories 'imagined' by this model, and updates the policy with on-policy GRPO (Group Relative Policy Optimization: sampling several trajectories for the same task and scoring them against each other), all without ever interacting with the real environment. Working in pixel space rather than a latent space keeps the imagined frames aligned with the visual features the VLA already learned from pretraining on internet images. The experiments use OpenVLA-OFT as the base policy, and WMPO outperforms GRPO and DPO baselines on both MimicGen simulation tasks and the real-robot Mobile ALOHA, with self-correcting behavior emerging along the way.
ExampleOn a real-robot 'insert the block onto the peg' task with a 5-millimeter clearance, the base policy succeeds 53% of the time, DPO reaches 60%, and WMPO reaches 70% (30 trials each).
- Also called
- World-Model-based Policy Optimization, WMPO: World Model-based Policy Optimization for Vision-Language-Action Models
- Related
- World Model · Reinforcement Fine-Tuning (RL Fine-Tuning) · Group Relative Policy Optimization · Learning in Imagination · OpenVLA-OFT · VLA-RFT
- Sources
- WMPO (arXiv:2511.09515)
WMPO 项目主页 (Chinese) - As of
- 2025-11