Mask R-CNN
AdvancedAn instance segmentation model from 2017, proposed by Kaiming He and colleagues, that outputs both detection boxes and a per-object mask.
Mask R-CNN was proposed in 2017 by Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick at Facebook AI Research, winning the Marr Prize (best paper) at ICCV 2017. Building on the two-stage detector Faster R-CNN, it adds a branch to each candidate region that predicts a pixel-level mask in parallel, so one network does object detection and instance segmentation (separating out which pixels belong to each object) at the same time; an additional branch can also do human keypoint detection. The paper introduces RoIAlign, which replaces RoIPool’s rounding with bilinear interpolation, removing the misalignment between features and pixels, which noticeably improves mask accuracy. The ResNet-101-FPN version reaches a mask AP of 35.7 on COCO at about 5 frames per second. In robotics, it was long the standard baseline for segmenting objects in grasping and picking pipelines, though it is now often replaced by or paired with open-vocabulary methods like the SAM family and Grounded-SAM.
ExampleIn a bin-picking pipeline, a Mask R-CNN fine-tuned on photos of the company’s own parts first segments a mask for every part in the bin, and the point cloud corresponding to each mask is then passed to the grasp-pose-detection module.
- Related
- Instance Segmentation · Object Detection · Mask · Segment Anything Model · Grounded SAM · Mean Average Precision
- Sources
- arXiv: Mask R-CNN
GitHub: facebookresearch/Detectron(注明 Marr Prize at ICCV 2017) (Chinese)