Learning in Imagination
想象中学习AdvancedLearning a world model first, then training the policy inside trajectories the model “imagines”.
Learning in imagination is a form of model-based reinforcement learning: first learn a world model from real interaction data — a model that predicts the next state and reward given the current state and action — then let the policy learn by trial and error inside trajectories the world model generates, touching the real environment as little as possible. David Ha and Jürgen Schmidhuber's 2018 paper “World Models” was an early demonstration of this idea. Danijar Hafner and colleagues' Dreamer series turned it into a general-purpose algorithm, imagining trajectories in a compressed latent space and backpropagating value-estimate gradients along the trajectory to train the policy. The upside is fewer real-world interactions and greater safety; the risk is that if the world model is inaccurate, the policy can learn tricks that only work “in the dream.”
ExampleDayDreamer (2022) ran Dreamer on a real robot: a quadruped learned to stand and walk from about an hour of real interaction. Dreamer 4 (2025), trained purely on offline data inside a world model, became the first such agent to obtain a diamond in Minecraft.
- Also called
- World-Model-Based RL, Training Inside a World Model, Dream Training
- Related
- World Model · Model-Based Reinforcement Learning · DreamerV3 · Dreamer 4 · DayDreamer · World Models
- Sources
- Dream to Control: Learning Behaviors by Latent Imagination (Dreamer, arXiv:1912.01603)
DayDreamer: World Models for Physical Robot Learning (arXiv:2206.14176)
Training Agents Inside of Scalable World Models (Dreamer 4, arXiv:2509.24527) - As of
- 2025-09