Embodied AI Glossary中文

DayDreamer

Advanced

Runs the Dreamer world model directly on real robots; a quadruped learns to walk from scratch in an hour.

DayDreamer was proposed in June 2022 by Pieter Abbeel and Ken Goldberg's groups at Berkeley, with authors including Danijar Hafner of the Dreamer series, published at CoRL 2022. The hard part of real-robot reinforcement learning is that trial and error is expensive, so most work trains in simulation first and transfers afterward. DayDreamer instead moves the Dreamer algorithm directly onto the real robot: as the robot interacts, it uses the collected data to learn a world model (predicting “what I'll see and what reward I'll get if I take this action”), and the policy learns mostly from imagined rollouts inside that world model, sharply reducing real-world trial and error. The same hyperparameters worked across four robots: an A1 quadruped learned to roll over, stand up, and walk from scratch in about 1 hour, and learned to resist being pushed in about 10 minutes; UR5 and xArm arms learned pick-and-place from camera images; a Sphero wheeled robot learned to navigate to a goal. It demonstrated the sample efficiency of world-model-based reinforcement learning in the real physical world.

ExampleWith no simulation pretraining and no manual resets, an A1 quadruped robot learned to roll over, stand up, and walk forward from scratch after about 1 hour of learning on real ground.

Also called
DayDreamer: World Models for Physical Robot Learning
Related
World Model · Learning in Imagination · Model-Based Reinforcement Learning · Real-World Reinforcement Learning · Recurrent State-Space Model · DreamerV3
Sources
DayDreamer (arXiv:2206.14176)
DayDreamer 项目主页(CoRL 2022) (Chinese)
As of
2022-06

See it in the full glossary →