DreamDojo
AdvancedA generalist robot world model NVIDIA released in 2026, pretrained on about 44,000 hours of human video.
DreamDojo is a generalist robot world model — a model that predicts future video frames conditioned on actions — released in February 2026, led by NVIDIA together with UC Berkeley, HKUST (Hong Kong University of Science and Technology), Stanford, and others. Robot data is scarce and expensive to collect, while videos of humans doing everyday things are abundant but carry no action labels. DreamDojo assembles about 44,700 hours of first-person human video (mainly its own crowdsourced dataset, DreamDojo-HV, about 43,800 hours, plus EgoDex and a small amount of lab data), which the paper describes as the largest video collection used to pretrain a world model at the time. It uses continuous latent actions learned from the video itself as a unified “proxy action.” The model first learns how objects behave under interaction from unlabeled video, and is then post-trained on robot data to turn it into a simulator that accepts real action commands. The model is built on Cosmos-Predict2.5, comes in 2B and 14B versions, and after distillation generates in real time at about 10.8 frames per second on an H100, making it usable for policy evaluation, real-time teleoperation, and model-based planning.
ExampleAfter post-training, it was adapted to robots including Fourier's GR-1, Unitree's G1, AgiBot, and the YAM arm: given a current frame and a sequence of actions, the model generates in real time what the video would look like after executing them, letting a policy be evaluated without running it on a real robot.
- Also called
- A Generalist Robot World Model from Large-Scale Human Videos
- Related
- World Model · Latent Action · Egocentric Video · EgoDex · NVIDIA Cosmos Predict · World-Model-based Policy Evaluation
- Sources
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos (arXiv 2602.06949)
DreamDojo 项目主页 (Chinese) - As of
- 2026-02