Recurrent State-Space Model
循环状态空间模型RSSMAdvancedThe core structure of the Dreamer world models, combining a deterministic recurrent state with a stochastic latent to predict the future.
The Recurrent State-Space Model is a latent dynamics model Danijar Hafner and colleagues (at Google Brain, DeepMind, and others) proposed in the 2019 PlaNet paper, and it later became the core of the Dreamer through DreamerV3 world models. Its state has two parts: a deterministic recurrent hidden state h, updated by a recurrent network such as a GRU, responsible for remembering history; and a stochastic latent z, representing the uncertain information at the current moment. Given the previous state and an action, the model first updates h, then predicts the next z; during training, a separate encoder infers z from the real image as a target, and the model is also required to reconstruct the image and predict the reward. PlaNet's comparisons show that a purely deterministic or purely stochastic version underperforms the combination of both. With it, an agent can 'imagine' many steps into the future inside latent space to plan or train a policy, greatly reducing how much it needs to interact with the real environment. DreamerV3 changes z into a discrete categorical distribution; 2025's Dreamer 4 switches to a Transformer-based world model instead.
ExampleDreamerV3 generates imagined trajectories with its RSSM world model and trains actor and critic networks on them, and the paper describes it as the first algorithm to mine diamonds from scratch in Minecraft with no human data.
- Also called
- RSSM
- Related
- World Model · Latent World Model · DreamerV3 · Dreamer 4 · Learning in Imagination · Recurrent Neural Network
- Sources
- Learning Latent Dynamics for Planning from Pixels (PlaNet, arXiv 1811.04551)
Mastering Diverse Domains through World Models (DreamerV3, arXiv 2301.04104)
Training Agents Inside of Scalable World Models (Dreamer 4, arXiv 2509.24527) - As of
- 2025-09