VLFM
VLFM(视觉语言前沿地图)AdvancedA method that scores exploration frontiers with a vision-language model to find a specified object in an unfamiliar environment, zero-shot.
VLFM was proposed by Naoki Yokoyama, Dhruv Batra, Bernadette Bucher, and colleagues at Georgia Tech and the Boston Dynamics AI Institute, posted to arXiv in December 2023 and published at ICRA 2024, where it won best paper in cognitive robotics. The task is object-goal navigation: finding an object of a given category in an environment the robot has never visited. It builds an occupancy map from depth data and identifies the frontier between known and unknown regions; at the same time, it uses BLIP-2 to compute the similarity between the current image and text prompts like 'the target is likely to be nearby,' recording this in a language value map, and then chooses the highest-value frontier to explore. Once the target is spotted, YOLOv7 or Grounding DINO detects it and Mobile-SAM segments it out for the robot to approach. The whole pipeline needs no navigation-specific training and achieved state-of-the-art results at the time by SPL (success weighted by path length) on Gibson, HM3D, and MP3D, and was deployed directly on a Boston Dynamics Spot.
ExampleIn an office building with no pre-built map, once Spot is given the category of object to find, it preferentially heads in the direction the vision-language model judges more likely to contain that object, and stops next to it once found.
- Also called
- Vision-Language Frontier Maps, VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation
- Related
- Object-Goal Navigation · Frontier-based Exploration · Zero-shot · Success weighted by Path Length · Occupancy Grid Map · VLMaps
- Sources
- VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation (arXiv 2312.03275)
VLFM 项目页 (Chinese) - As of
- 2024-05