PointNet / PointNet++
PointNetCommonA pioneering network that classifies and segments point clouds directly, with PointNet++ its hierarchical upgrade.
PointNet was proposed by Charles R. Qi, Leonidas Guibas, and colleagues at Stanford, published at CVPR 2017. Before it, point clouds were usually converted to voxels or multi-view images before processing, which lost detail. PointNet instead extracts a feature for each point with a shared multi-layer perceptron (MLP), then aggregates these into a global feature with max pooling; because max pooling is order-independent, shuffling the points doesn't change the result (permutation invariance), and the network can directly perform classification, part segmentation, and scene semantic segmentation. Its shortcoming is that it doesn't model local neighborhoods. PointNet++ (NeurIPS 2017) fixes this: it first picks center points with farthest point sampling, recursively applies PointNet within their neighborhoods, and extracts multi-scale local features layer by layer, while also adapting to uneven point density. Both remain common baselines for point cloud encoders today.
Example3D Diffusion Policy (DP3)'s point cloud encoder follows the ‘per-point MLP plus max pooling’ idea, using only three MLP layers; the paper found that more complex encoders like PointNet and PointNet++ actually underperformed this small encoder on its manipulation tasks.
- Also called
- PointNet++
- Related
- Point Cloud · Point Cloud Encoder · Farthest Point Sampling · Point Cloud Segmentation · Point Transformer V3 · 3D Diffusion Policy
- Sources
- PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (arXiv 1612.00593)
PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space (arXiv 1706.02413)
3D Diffusion Policy (arXiv 2403.03954)