Cosmos 3
AdvancedNVIDIA's third-generation omnimodal world model: a single model that understands video, generates it, and also outputs actions.
This is the third generation of NVIDIA's Cosmos world foundation model family, released in May 2026, with the technical report posted to arXiv in June 2026. The first two generations split understanding (Cosmos Reason) and video generation (Cosmos Predict) into separate models; Cosmos 3 merges them into one using a mixture-of-transformers (MoT) architecture: an autoregressive language tower reads text and reasons, a diffusion generation tower outputs images, video, audio, and action sequences, and the two towers share attention. The same set of weights can act as a vision-language model, a text-to-video or image-to-video model, a world simulator that predicts future frames from actions, and can also be post-trained into a robot policy. It's released in Super (64B) and Nano (16B) sizes, with an Edge (4B) version for Jetson edge devices added in July 2026; code and weights use the OpenMDW-1.1 license.
ExampleAccording to the technical report, Cosmos3-Nano-Policy, post-trained on DROID data, ranked first on the RoboArena real-robot leaderboard (a snapshot from May 30, 2026) with an Elo of 1870.
- Also called
- Cosmos3, NVIDIA Cosmos 3, Cosmos 3: Omnimodal World Models for Physical AI
- Related
- NVIDIA Cosmos · World Foundation Model · Mixture-of-Transformers · Unified Multimodal Model · World Action Model · Cosmos Policy
- Sources
- Cosmos 3: Omnimodal World Models for Physical AI (arXiv:2606.02800)
NVIDIA Cosmos GitHub(Cosmos 3 模型家族与发布记录) (Chinese) - As of
- 2026-07