Embodied AI Glossary中文

NVIDIA Cosmos Predict

Cosmos PredictAdvanced

The video world model branch of NVIDIA's Cosmos family, generating what comes next from text, an image, or video.

Cosmos Predict is the branch of NVIDIA's Cosmos world foundation model family responsible for predicting future frames. Predict1 launched with the Cosmos platform in January 2025, with 7B/14B diffusion models and several autoregressive model sizes; Predict2 opened up in June 2025, with 2B and 14B video models; Predict2.5, from October 2025, switched to flow-based generation, uses Cosmos-Reason1 as its text encoder, and merges text-to-video, image-to-video, and video continuation into one model. It's positioned as a fine-tunable base for robotics and autonomous driving: post-trained on robot data, it can become an action-conditioned simulator, a multi-view generator, or be used directly as a policy (as in Cosmos Policy). Starting in 2026, NVIDIA has shifted its main focus to the unified Cosmos 3, and the Predict repository is no longer under active development.

ExampleNVIDIA provides a version of Cosmos-Predict2 post-trained with action conditioning on the Bridge robot dataset: given the current frame and a sequence of arm actions, it generates video of what happens after executing them; GR00T Dreams also uses a post-trained version of it to generate robot training video.

Also called
Cosmos-Predict1, Cosmos-Predict2, Cosmos-Predict2.5
Related
NVIDIA Cosmos · World Foundation Model · Video Generation Model · Cosmos Policy · DreamGen · Cosmos 3
Sources
nvidia-cosmos/cosmos-predict2.5 GitHub 仓库 (Chinese)
nvidia-cosmos/cosmos-predict2 GitHub 仓库 (Chinese)
Cosmos World Foundation Model Platform for Physical AI (arXiv 2501.03575)
As of
2026-06

See it in the full glossary →