Embodied AI Glossary中文

Cosmos Policy

Advanced

Fine-tunes a video generation model directly into a robot policy that predicts actions, future frames, and success confidence together.

Cosmos Policy was proposed by NVIDIA and Stanford in January 2026, first author Moo Jin Kim (also first author of OpenVLA-OFT), and accepted at ICLR 2026. It starts from the video generation model Cosmos-Predict2-2B and, without changing its architecture or adding an action head, post-trains it just once on robot demonstration data: proprioceptive state, action chunks, and value (expected return) are all encoded as “latent frames” and inserted into the video diffusion model's latent sequence, denoised together with the image frames. This way, one pass gives three things at once: the action to execute, the future frames that will result, and how confident the model is that this step succeeds. At inference, it can execute the action directly, or sample several candidates and use the predicted future and value to pick the best one (test-time planning). It reaches 98.5% average success on LIBERO and 67.1% on RoboCasa (50 demonstrations per task), representative of the “fine-tune a video model directly into a policy” approach.

ExampleOn real ALOHA bimanual tasks such as putting candy into a sealed bag, in planning mode Cosmos Policy first samples 8 candidate action chunks, then scores them using its own predicted future frames and value, and executes whichever one it's most confident in.

Also called
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
Related
NVIDIA Cosmos Predict · Video Generation Model · World Action Model · Value Function · OpenVLA-OFT · LIBERO Benchmark
Sources
Cosmos Policy (arXiv:2601.16163)
Cosmos Policy 项目主页(NVIDIA Research) (Chinese)
As of
2026-01

See it in the full glossary →