DreamGen
DreamGen(GR00T Dreams)AdvancedAn NVIDIA 2025 method that uses a video world model to generate robot videos and infer actions from them to synthesize training data.
DreamGen is a synthetic-data pipeline proposed by NVIDIA's GEAR Lab and others in May 2025; its open-source implementation is called the Isaac GR00T-Dreams blueprint, announced that same month at Computex in Taipei. Teaching a robot a new skill usually requires a human to teleoperate it to collect large amounts of data. DreamGen works in four steps: first fine-tune an image-to-video generation model on data from the target robot (the open-source implementation uses Cosmos-Predict2); then, given a single starting frame and a language instruction, have the model generate a realistic video of the robot performing the new action in a new environment; next, use an inverse dynamics model or a latent action model to infer pseudo action labels from that video; and finally train a visuomotor policy on these “neural trajectories” together with real data. The paper also proposes DreamGen Bench, finding that higher video-generation quality leads to a better-trained policy. NVIDIA says that using GR00T-Dreams, it generated training data for GR00T N1.5 in 36 hours, versus roughly three months for manual collection.
ExampleUsing only teleoperation data from a single pick-and-place task in one environment, plus videos generated by DreamGen, the GR1 humanoid robot learned 22 new behaviors and could perform them in 10 environments it had never seen.
- Also called
- GR00T Dreams, GR00T-Dreams, Isaac GR00T-Dreams, Unlocking Generalization in Robot Learning through Video World Models
- Related
- Neural Trajectories · Video Generation Model · Inverse Dynamics Model · Latent Action Model · NVIDIA Isaac GR00T N1 · NVIDIA Cosmos Predict
- Sources
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models (arXiv 2505.12705)
DreamGen 项目主页(NVIDIA GEAR) (Chinese)
NVIDIA Computex 2025 新闻稿:GR00T N1.5 与 GR00T-Dreams (Chinese) - As of
- 2025-05