Embodied AI Glossary中文

VLMaps

Advanced

A system that writes vision-language features into a 3D map, letting a robot navigate using phrases like 'three meters to the right of the chair.'

VLMaps was proposed by Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard at the University of Freiburg, Google Research, and the Nuremberg Institute of Technology, posted to arXiv in October 2022 and published at ICRA 2023. As it moves through an environment, the robot builds a map from RGB-D video: an open-vocabulary segmentation model, LSeg, computes a vision-language feature for every pixel, which is projected onto 3D surfaces using depth and pose and then compressed into a top-down grid map. At query time, words like 'sofa' or 'fridge' are encoded with a text encoder and compared against the map's features by similarity, letting the system locate any object — an open-vocabulary semantic map. Complex instructions are first handed to a large language model (the GPT-3 family), which writes code that calls the map's interface, so the system can handle spatial relations like 'between two objects' or 'three meters to the right.' The same map can also generate a separate obstacle map for each different robot. VLMaps is an early representative of connecting foundation models to navigation maps.

ExampleTold to 'move to the spot three meters to the right of the chair,' the language model turns the instruction into code that first looks up the chair's position on the map, then computes the point three meters to its right for the robot to head toward.

Also called
Visual Language Maps, Visual Language Maps for Robot Navigation
Related
Semantic Map · Vision-and-Language Navigation · Open-vocabulary · Code as Policies · CLIP · VLFM
Sources
Visual Language Maps for Robot Navigation (arXiv 2210.05714)
VLMaps 项目主页 (Chinese)
vlmaps/vlmaps (GitHub)
As of
2023-03

See it in the full glossary →