Visual Grounding
视觉定位(Grounding)CommonFinds the location in an image that a phrase or sentence refers to.
Visual grounding links language to image regions: given an image and a piece of text, it outputs the box, mask, or point of the object referred to. Finding a region for a phrase is phrase grounding; finding the single region that a uniquely identifying description like “the red cup on the left” points to is referring expression comprehension (REC), commonly evaluated on the RefCOCO family of datasets. It demands more than plain detection — understanding attributes and spatial relationships too. In 2023, IDEA Research’s Grounding DINO could locate objects from arbitrary text, and today most vision-language models can output boxes or point coordinates directly. When a robot follows an instruction to pick something up, grounding is usually the first step. Note that the Chinese term for this, 视觉定位, is also commonly used for a robot figuring out its own location from a camera (visual localization) — a different, unrelated task despite the similar name.
ExampleGiven the instruction “hand me the red cup on the left,” the model first draws a box around that one cup in the camera image, ignoring the blue cup next to it, then projects the depth points inside the box into 3D space for the grasping module to use.
- Also called
- Referring Expression Comprehension, REC, Phrase Grounding
- Related
- Referring Expression Segmentation · 3D Visual Grounding · Grounding DINO · Open-Vocabulary Object Detection · Pointing · Vision-Language Model
- Sources
- Modeling Context in Referring Expressions (RefCOCO / RefCOCO+ / RefCOCOg, ECCV 2016)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection