Embodied AI Glossary中文

Point Cloud Encoder

点云编码器Advanced

A module that converts an unordered set of 3D points into feature vectors a neural network can use.

A point cloud encoder turns a point cloud from a depth camera or lidar — an unordered set of 3D points with xyz coordinates, and sometimes color — into features. The difficulty is that points have no fixed order and no fixed count, so the convolutions built for image grids don't directly apply. Stanford's 2016 PointNet extracts per-point features with a shared MLP, then aggregates them with an order-independent operation like max pooling, pioneering the direct-on-point-cloud approach; stronger architectures followed, such as PointNet++ and the Point Transformer series. Any policy that uses 3D input in embodied AI depends on this: 3D Diffusion Policy (DP3) first downsamples the point cloud to 512 or 1024 points with farthest point sampling, then uses a lightweight encoder of three MLP layers plus max pooling to get a 64-dimensional feature, and the paper's ablations show this outperforms more complex encoders like PointNet++.

ExampleDP3's point cloud encoder deliberately skips the color channel and uses only geometric coordinates, which the paper says generalizes better to changes in an object's appearance.

Also called
Point Cloud Backbone
Related
Point Cloud · PointNet / PointNet++ · Point Transformer V3 · 3D Diffusion Policy · Farthest Point Sampling · Vision Encoder
Sources
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (arXiv:1612.00593)
3D Diffusion Policy (arXiv:2403.03954)

See it in the full glossary →