Embodied AI Glossary中文

Open-Vocabulary Segmentation

开放词汇分割Advanced

Segmenting exactly the pixels that match any text description of a category, even one the model never saw during training.

Traditional segmentation models can only separate out the few dozen to few hundred categories fixed at training time. Open-vocabulary segmentation instead lets a category be specified at test time with arbitrary text — for example, “the blue dish sponge” — and the model outputs a pixel mask for it. It is built on CLIP-style image-text contrastive pretraining: features for each pixel or region of the image and features for the text are placed in the same embedding space, and similarity determines the match. Representative work includes LSeg (ICLR 2022); later work combined open-vocabulary detectors with SAM in Grounded-SAM, and SAM 3 accepts noun-phrase prompts directly. For robots, this means finding a particular item on a table no longer requires labeling new data and retraining for every new object.

ExampleA user says “put the charging cable away”; the robot feeds the words “charging cable” into a segmentation model as text, gets back a pixel mask for the cable, and combines it with a depth map to compute a grasp point.

Also called
Open-Set Segmentation, Open-Vocabulary Semantic Segmentation
Related
Open-Vocabulary Object Detection · Semantic Segmentation · CLIP · Grounded SAM · SAM 3 · Open-vocabulary
Sources
Language-driven Semantic Segmentation (LSeg, arXiv 2201.03546)
Towards Open Vocabulary Learning: A Survey (arXiv 2306.15880)

See it in the full glossary →