Semantic Map
语义地图AdvancedA map that labels what each place is, on top of its geometry, so it can be queried by object name or natural language.
An ordinary SLAM map only records where obstacles are and where surfaces sit; a semantic map goes further, attaching category labels or features to the points, cells, or objects in the map, so it can answer questions like “where is the refrigerator” or “where is the kitchen.” Early approaches labeled the map using detection or segmentation networks with a fixed set of categories; in recent years, open-vocabulary semantic maps have become popular, storing features from vision-language models such as CLIP or LSeg directly in the map so it can be queried with arbitrary natural language. A representative example is VLMaps (2022, from the University of Freiburg, Google, and others), which extracts pixel-level vision-language features from RGB-D video, back-projects them into 3D using depth, and stores them in a top-down grid map, letting it follow navigation instructions with spatial relationships such as “go between the sofa and the TV.” Semantic maps can also be organized as 3D scene graphs, and are a common intermediate representation for object-goal navigation, vision-and-language navigation, and mobile manipulation.
ExampleVLMaps lets a LoCoBot and a drone share the same language map, navigating from instructions such as “move three meters to the right of the chair.”
- Also called
- Open-Vocabulary Semantic Map, Language Map
- Related
- Semantic SLAM · VLMaps · ConceptGraphs · 3D Scene Graph · Object-Goal Navigation · Occupancy Grid Map
- Sources
- VLMaps 项目主页 (Chinese)
arXiv 2210.05714: Visual Language Maps for Robot Navigation