World Foundation Model
世界基础模型WFMCommonA general-purpose world model pretrained on huge amounts of real video that can be fine-tuned into specialized world models.
The term World Foundation Model was introduced in NVIDIA's January 2025 paper announcing the Cosmos platform, describing a general-purpose world model that can be post-trained into customized world models for specific applications. A world model's job is to predict what the world will look like next, given the current view plus an action or instruction. WFM borrows the large language model playbook — pretrain at massive scale first, then fine-tune per task: Cosmos filtered about 100 million clips out of roughly 20 million hours of raw video and used them to train both diffusion and autoregressive model variants. Its uses include generating synthetic training data for robots and self-driving cars, training and evaluating policies inside the model, and turning simulated footage into photorealistic footage. NVIDIA has since released the Cosmos Predict, Transfer, and Reason series, as well as Cosmos 3, which puts reasoning, video, and action generation into a single model.
ExampleThe Cosmos paper post-trains a pretrained WFM into a robot-manipulation version: given the current view and a sequence of robot actions, it predicts the video of what happens after executing them.
- Also called
- WFM
- Related
- World Model · Foundation Model · NVIDIA Cosmos · Video Generation Model · Synthetic Data · Physical AI
- Sources
- Cosmos World Foundation Model Platform for Physical AI (arXiv 2501.03575)
NVIDIA Glossary: What Are World Models?
NVIDIA Cosmos - As of
- 2026-09