DINO-WM
AdvancedA world model that predicts the future in DINOv2's image-feature space, able to plan zero-shot toward new goals after training.
DINO-WM was proposed by Gaoyue Zhou and Hengkai Pan at NYU with Yann LeCun and Lerrel Pinto (the former also at Meta FAIR), released in November 2024 and published at ICML 2025. Many world models need to reconstruct pixels, or are tied to a specific task reward. DINO-WM instead uses a pretrained, frozen DINOv2 to extract patch features, and trains only a predictor: given the current features and an action, it predicts the next step's features, with no image reconstruction at all. Training uses only offline collected trajectories, needing no expert demonstrations, reward model, or inverse-dynamics model. At test time, given a goal image, it uses model-predictive control (MPC) to search for a sequence of actions whose predicted future features come as close as possible to the goal's features — zero-shot planning. It represents the approach of building a world model inside a pretrained representation's latent space, in a similar spirit to JEPA and V-JEPA 2.
ExampleIn the Push-T task, given a goal image of the desired arrangement, DINO-WM rolls out different pushing actions in feature space and picks the sequence that brings the T-shaped block closest to the goal pose; the same method is also used for maze navigation and manipulating rope and granular materials.
- Also called
- DINO World Model, DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
- Related
- World Model · Latent World Model · DINOv2 · Model Predictive Control · Joint-Embedding Predictive Architecture · V-JEPA 2
- Sources
- DINO-WM (arXiv:2411.04983)
DINO-WM 项目主页 (Chinese)
Gaoyue Zhou 个人主页(标注 ICML 2025) (Chinese) - As of
- 2025