Embodied AI Glossary中文

Ctrl-World

Advanced

A manipulation world model that generates multi-view future frames from actions, used to evaluate and improve policies “in imagination.”

Ctrl-World was proposed in October 2025 by Chelsea Finn's group at Stanford with Jianyu Chen's group at Tsinghua, accepted at ICLR 2026. Evaluating and improving a general-purpose robot policy usually needs a large number of real-robot trials, which is slow and expensive. Ctrl-World is initialized from the 1.5-billion-parameter video diffusion model Stable Video Diffusion, and trained on the DROID dataset (about 95,000 trajectories across 564 scenes) into a world model that generates future frames given an action, with three modifications: jointly predicting two third-person views plus a wrist camera; sparsely sampling history frames and embedding the arm's pose so the model can retrieve relevant past frames by pose and stay consistent over long horizons; and injecting actions frame by frame for centimeter-level control precision. A policy can interact inside it continuously for more than 20 seconds; it's used to rank π0, π0-FAST, and π0.5, matching real-robot results, and fine-tuning π0.5 on successful trajectories picked from imagined rollouts raised its success rate on unfamiliar instructions and objects from 38.7% to 83.4%.

ExampleGiven π0.5 an instruction it hasn't practiced, running 400 imagined trials in Ctrl-World with reworded instructions and randomized arm resets, then manually picking 25–50 successful trajectories and fine-tuning the policy on them for 2,000 steps, is enough to meaningfully improve it.

Also called
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
Related
World Model · World-Model-based Policy Evaluation · DROID (Distributed Robot Interaction Dataset) · π0.5 · Stable Video Diffusion · Synthetic Data
Sources
Ctrl-World (arXiv:2510.10125)
Ctrl-World 项目主页(ICLR 2026) (Chinese)
As of
2026-03

See it in the full glossary →