Embodied AI Glossary中文

DreamerV3

Common

A world-model reinforcement-learning algorithm that trains a policy by imagining rollouts, using one fixed set of hyperparameters.

The Dreamer series is led by Danijar Hafner. The original 2019 Dreamer proposed learning behavior by “imagining” the future inside a latent space; 2020's DreamerV2 switched to discrete latent variables and was the first world-model-based agent to reach human-level performance on Atari, beating top single-GPU model-free methods like Rainbow and IQN; DreamerV3 was posted publicly in January 2023 and published in Nature in 2025. The approach: first learn a world model from interaction experience that predicts the next latent state and reward, then train the policy on imagined trajectories generated by that model, using the real environment mainly to collect data. Through robustness tricks like normalization, balancing, and transformations, it solves more than 150 different tasks with the same fixed hyperparameters, and it became the first algorithm to mine diamonds in Minecraft from scratch without human data or a curriculum.

ExampleIn Minecraft, DreamerV3 learned from scratch, without human demonstrations, to chop trees, craft tools, and eventually mine a diamond.

Also called
Dreamer Series, DreamerV1, DreamerV2
Related
World Model · Recurrent State-Space Model · Learning in Imagination · Model-Based Reinforcement Learning · DayDreamer · Dreamer 4
Sources
DreamerV3 (arXiv 2301.04104)
DreamerV3 项目页(Danijar Hafner) (Chinese)
DreamerV2: Mastering Atari with Discrete World Models (arXiv 2010.02193)
As of
2025

See it in the full glossary →