Embodied AI Glossary中文

Instance Segmentation

实例分割Common

Cutting out every object in an image pixel by pixel, keeping separate objects of the same class distinct from each other.

Instance segmentation has to answer two questions at once — ‘what is this’ and ‘which one is this’ — generating a separate pixel-level mask and class label for every countable object (‘things,’ in the terminology of the field) in an image. This differs from semantic segmentation, which only assigns a class to each pixel: three cups on a table would merge into a single ‘cup’ region under semantic segmentation, whereas instance segmentation splits them into cup 1, cup 2, and cup 3. The representative method is Mask R-CNN, from Kaiming He and colleagues in 2017, which adds a parallel mask-prediction branch on top of an object-detection box. Panoptic segmentation, proposed in 2018, further folds in uncountable background regions (‘stuff’) such as sky and ground. Robot grasping usually uses instance segmentation first to cut out the target object, then crops the corresponding point cloud to compute a grasp pose.

ExampleTwo identical apples sit side by side on a table. Semantic segmentation only gives one merged ‘apple’ region, while instance segmentation outputs two separate masks, letting the robot grasp specifically the one on the left as instructed.

Also called
Instance-Level Segmentation
Related
Semantic Segmentation · Panoptic Segmentation · Mask · Object Detection · Mask R-CNN · Segment Anything Model
Sources
Mask R-CNN (arXiv:1703.06870)
Panoptic Segmentation (arXiv:1801.00868)

See it in the full glossary →