Embodied AI Glossary中文

LM-Nav

Advanced

Combines GPT-3, CLIP, and a visual navigation model so a robot can navigate by following natural-language instructions.

LM-Nav was released in July 2022 by Dhruv Shah, Brian Ichter, Sergey Levine, and colleagues, published at CoRL 2022. Directing a robot's navigation with language usually needs a lot of trajectory data with text descriptions, which is expensive to annotate. LM-Nav does no fine-tuning at all and uses no language-labeled robot data; instead, it combines three off-the-shelf pretrained models: the large language model GPT-3 breaks the instruction down into a sequence of landmarks; the image-text model CLIP judges which landmark corresponds to what the robot's camera sees; and the visual navigation model ViNG builds a topological map of the environment from previously collected images and runs a go-to-point policy. The system then searches for the shortest route that passes through those landmarks in order. It completed long-distance navigation in real outdoor environments, and is an early representative example of assembling a robot system out of foundation models.

ExampleFor example, the user says “go past the stop sign and head to the white building”; GPT-3 extracts the two landmarks “stop sign” and “white building,” CLIP locates the corresponding positions in the topological map, and ViNG drives to them in sequence.

Also called
Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
Related
Vision-and-Language Navigation · CLIP · Topological Map · LLM-based Task Planning · GNM · ViNT
Sources
arXiv 2207.04429: LM-Nav
LM-Nav 项目主页 (Chinese)
As of
2022-07

See it in the full glossary →