Embodied AI Glossary中文

Point Cloud Segmentation

点云分割Advanced

Labeling every point in a point cloud with which object category, or which individual object, it belongs to.

Point cloud segmentation is the 3D counterpart of 2D image segmentation: semantic segmentation assigns each point a category (table, cup, floor); instance segmentation further distinguishes separate individuals within the same category (cup 1, cup 2); and part segmentation breaks a single object down into parts such as a handle or a lid. PointNet, from Stanford in 2017, was the first network to operate directly on unordered point sets for classification and segmentation, and it was followed by backbones such as PointNet++ and the Point Transformer series. Common benchmarks include the indoor datasets ScanNet and S3DIS, and the outdoor dataset SemanticKITTI. A robot arm typically segments out the target object’s points before estimating its pose or detecting a grasp; a mobile robot uses it to separate the ground, obstacles, and traversable area.

ExampleIn a point cloud of a tabletop scene, the points belonging to “cup” are segmented out on their own and fed into a grasp-detection network.

Also called
3D Semantic Segmentation, 3D Instance Segmentation, Point Cloud Semantic Segmentation
Related
Point Cloud · Semantic Segmentation · Instance Segmentation · PointNet / PointNet++ · Point Transformer V3 · Grasp Pose Detection
Sources
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (arXiv 1612.00593)
Point Transformer V3: Simpler, Faster, Stronger (arXiv 2312.10035)

See it in the full glossary →