Grounded SAM
Grounded-SAMAdvancedAn open-source pipeline that boxes and finely segments the object a text phrase refers to in an image.
Grounded SAM is an open-world vision pipeline from IDEA Research, with a technical report released in January 2024. Its core idea is chaining two models together: Grounding DINO finds an object from a text description and outputs a detection box, and SAM (the Segment Anything Model) takes that box and outputs a pixel-level mask. This means any text — say, “red cup” — can produce a mask for the matching object, with no retraining needed for a new category; it can also connect to RAM and BLIP for automatic labeling, and to Stable Diffusion for image editing. The report states it reaches 48.7 mAP on the SegInW zero-shot segmentation benchmark. In robotics, it’s commonly used to specify a target by language and extract that object’s point cloud for grasping or pose estimation, and also for automatic data labeling; the later Grounded SAM 2 connects to SAM 2 and can track objects across video.
ExampleGiven the instruction “put the banana in the bowl,” a robot first uses Grounded-SAM to segment masks for “banana” and “bowl” separately, then combines these with a depth map to compute both objects’ 3D positions, handing this to the grasping and motion-planning modules.
- Also called
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks, Grounded-Segment-Anything
- Related
- Grounding DINO · Segment Anything Model · Open-Vocabulary Segmentation · Open-Vocabulary Object Detection · SAM 2 · Auto-labeling
- Sources
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks (arXiv 2401.14159)
IDEA-Research/Grounded-Segment-Anything GitHub - As of
- 2024-01