NaVid
AdvancedA 2024 video-LLM navigation method from Peking University and others that decides the next step using only monocular video.
NaVid was proposed in February 2024 by He Wang's group at Peking University together with the Beijing Academy of Artificial Intelligence, the University of Adelaide, Galbot, and others, published at RSS 2024, addressing vision-and-language navigation (reaching a destination by following a one-sentence instruction). Earlier methods mostly relied on maps, odometry, or depth maps; NaVid uses only monocular RGB video: the current frame is compressed to 64 tokens and each history frame to 4, fed into a video large language model built on Vicuna-7B, which outputs the next step directly in text, such as how far to move forward, how much to turn, or to stop. Using only RGB, it reaches a 37.4% success rate on R2R-CE, and transfers to a real robot. The follow-up Uni-NaVid (RSS 2025) merges four task types — instruction navigation, object finding, embodied question answering, and person following — into a single model, trained on 3.6 million samples.
ExampleTested across 4 real indoor scenes with 200 instructions total, NaVid running on a Turtlebot4 achieved about 66% success on simple instructions and about 48% on multi-step compound instructions.
- Also called
- Uni-NaVid, Video-based VLM Plans the Next Step for Vision-and-Language Navigation
- Related
- Vision-and-Language Navigation · Room-to-Room · Vision-Language Model · NaVILA · NavFoM (Galbot) · Sim-to-Real Transfer
- Sources
- NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation (arXiv 2402.15852)
NaVid 项目页 (Chinese)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks (arXiv 2412.06224) - As of
- 2025-02