NVIDIA Cosmos Reason
Cosmos ReasonAdvancedNVIDIA's reasoning vision-language model for physical AI, which watches a video, reasons through it in writing, then answers.
Cosmos Reason is the vision-language model in NVIDIA's Cosmos family responsible for understanding and reasoning. Cosmos-Reason1, from March 2025, trains in two steps: supervised fine-tuning on physical-AI data first, then reinforcement learning, so the model writes out a chain of thought before answering; the paper covers 7B and 56B sizes, with the open-sourced version being the 7B built on Qwen2.5-VL. It focuses on physical common sense (space, time, basic physics) and embodied reasoning (what a robot should do next). Cosmos-Reason2, from December 2025, switched to Qwen3-VL, offered at 2B and 8B, with a 32B added in April 2026. Common uses include flagging physical errors in generated video, filtering and annotating training data, and serving as a robot's planning module; Cosmos-Predict2.5 uses it as a text encoder, and the driving model Alpamayo also uses it as its backbone.
ExampleGiven an AI-generated robot manipulation video and asked “does this footage obey physics,” it first outputs its reasoning process (such as whether an object moves without cause, or clips through another object), then gives its judgment, which can be used to automatically filter out unqualified synthetic data.
- Also called
- Cosmos-Reason1, Cosmos-Reason2
- Related
- Vision-Language Model · Embodied Reasoning Model · Chain-of-Thought · NVIDIA Cosmos · Qwen-VL · NVIDIA Alpamayo
- Sources
- Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning (arXiv 2503.15558)
nvidia-cosmos/cosmos-reason2 GitHub 仓库 (Chinese)
nvidia-cosmos/cosmos-reason1 GitHub 仓库 (Chinese) - As of
- 2026-04