Latent World Model
隐空间世界模型AdvancedA world model that predicts 'what happens if this action is taken' in a compressed latent state instead of pixels.
A world model predicts the future given the current state and an action. Predicting future images directly means generating a huge amount of pixel-level detail irrelevant to decision-making, which is slow and hard; a latent world model instead encodes observations into a low-dimensional latent state first, predicts only the next latent state within that latent space, and does planning or policy training there too. Hafner and colleagues' 2018 PlaNet introduced the Recurrent State-Space Model (RSSM), which mixes deterministic and stochastic components, and the later Dreamer series uses it to train policies 'in imagination.' A different line of work skips reconstructing pixels altogether: DINO-WM directly predicts the patch features of the pretrained vision model DINOv2; Meta's 2025 V-JEPA 2-AC trains an action-conditioned predictor on fewer than 62 hours of DROID robot video and achieves zero-shot pick-and-place on Franka arms in two different labs. The difficulty is that the latent state can't be inspected directly, and methods that skip pixel reconstruction also have to guard against representation collapse (all inputs getting encoded into nearly identical vectors).
ExampleFor pick-and-place, V-JEPA 2-AC encodes the goal image into a latent vector, predicts the outcome of several candidate actions inside latent space, and executes whichever one's predicted result lands closest to the goal — without ever generating an image.
- Also called
- Latent Dynamics Model
- Related
- World Model · Recurrent State-Space Model · DINO-WM · V-JEPA 2 · Joint-Embedding Predictive Architecture · Learning in Imagination
- Sources
- Learning Latent Dynamics for Planning from Pixels (PlaNet, arXiv:1811.04551)
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning (arXiv:2411.04983)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv:2506.09985) - As of
- 2025-06