Embodied AI Glossary中文

Referring Expression Segmentation

指代表达分割RESAdvanced

Given a sentence describing one specific object, precisely segmenting out exactly that object in the image.

Referring expression segmentation takes an image plus a natural-language description (such as “the two people sitting on the bench on the right”) and outputs a pixel-level mask of the object the sentence refers to. It is more fine-grained than open-vocabulary segmentation: the latter segments every object of a named category, while referring segmentation must pick out one specific instance based on color, position, relationships, and other descriptive cues. Hu et al. proposed an early end-to-end method in 2016, encoding the sentence with an LSTM and fusing it with a convolutional network’s feature map to predict a mask pixel by pixel. Common benchmarks include RefCOCO, RefCOCO+, and G-Ref. Today the task is often handled by multimodal large models or grounding models paired with a SAM-style segmentation model. When a robot hears “hand me that red cup on the left,” it must first use this to find the pixel region of the target, then combine that with depth to get a 3D position for grasping.

ExampleGiven the instruction “pick up the blue block on the far left of the table,” the model outputs a mask for exactly that block in the wrist camera’s image.

Also called
RES, Referring Image Segmentation
Related
Visual Grounding · Open-Vocabulary Segmentation · Instance Segmentation · Mask · Grounded SAM · Language Grounding
Sources
Segmentation from Natural Language Expressions (arXiv 1603.06180)

See it in the full glossary →