Embodied AI Glossary中文

Navigation World Models (Meta)

导航世界模型NWMAdvanced

A late-2024 Meta video world model for navigation that imagines what you'd see after taking a given action.

Navigation World Models was proposed by Amir Bar, Yann LeCun, and colleagues at Meta FAIR, together with NYU and Berkeley, posted to arXiv in December 2024, and received a CVPR 2025 best-paper honorable mention. It is a controllable video-generation model: given past frames and a navigation action (which way to go, how much to turn), it predicts the footage you'd see next. The model is a 1-billion-parameter conditional diffusion Transformer (CDiT), trained on first-person video from both humans and robots. With it, a robot can simulate several candidate routes inside the model first and pick whichever reaches the goal; it can also score and rank trajectories sampled by an existing policy like NoMaD; and it can imagine walking through an unfamiliar environment given just a single photo. The authors also note that generating for a long time in an unfamiliar environment gradually drifts the footage back toward the training data.

ExampleGiven a starting photo and a goal photo, NWM generates the along-the-way footage for a batch of candidate action sequences one by one, and picks whichever ends up closest to the goal photo to execute.

Also called
NWM
Related
World Model · Video Prediction Model · Diffusion Transformer · NoMaD · Image-Goal Navigation · Meta Fundamental AI Research
Sources
Navigation World Models (arXiv 2412.03572)
Navigation World Models 项目页 (Chinese)
As of
2025-06

See it in the full glossary →