Visual Prompting (Set-of-Mark)
视觉提示(Set-of-Mark 标记提示)SoMAdvancedDrawing numbers, boxes, or dots on an image so a vision-language model can point at a location just by naming its label.
Visual prompting means guiding a vision-language model (VLM) by marking up the input image itself, without changing any model weights. The flagship method is Set-of-Mark (SoM), introduced by Microsoft Research in October 2023: a segmentation model such as SAM or SEEM first cuts the image into regions, each region gets a number, mask, or box overlaid on it, and the marked-up image is then shown to a model like GPT-4V. It's hard for a VLM to output precise pixel coordinates directly, but answering “region 3” is much easier — SoM even beat specially fine-tuned models on zero-shot RefCOCOg referring segmentation. In robotics, this is often used to turn “where to grasp, where to place it” into a multiple-choice question: MOKA (2024, UC Berkeley) marks candidate keypoints and waypoints on the image and has the VLM pick the grasp point and motion path, which then gets converted into arm actions; PIVOT instead draws a batch of candidate actions on the image and has the VLM select and iteratively narrow them down.
ExampleA tabletop photo is first segmented with SAM into individual objects labeled 1, 2, 3, and so on. Asking GPT-4V “which one is the red cup” gets back the answer “2,” and the program can then pull the precise mask for region 2 and hand it to the grasping module.
- Also called
- SoM, Set-of-Mark Prompting, Mark-based Visual Prompting
- Related
- Vision-Language Model · Visual Grounding · Segment Anything Model · MOKA · PIVOT · Prompt / Prompt Engineering
- Sources
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V (arXiv:2310.11441)
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting (arXiv:2403.03174)