Point Cloud
点云EssentialData made of many 3D coordinate points used to represent the shape of an object or scene.
A point cloud is a set of discrete points in 3D space, each carrying X, Y, Z coordinates and optionally color, a surface normal, a timestamp, and so on. It typically comes from lidar, from a depth camera (a depth map back-projected using the camera's intrinsics), or from multi-view 3D reconstruction. Point clouds are unordered, sparse, and variable in size, so ordinary image convolutional networks don't directly apply to them, which is why specialized point cloud networks like PointNet exist. Common processing steps include voxel downsampling or farthest point sampling (both reduce the point count), aligning two clouds with ICP (Iterative Closest Point, a registration method), and segmenting out a target object; the open-source PCL library defines the widely used PCD file format. In embodied AI, 3D Diffusion Policy (DP3) uses a single-view point cloud as its policy input, and grasp-detection methods such as AnyGrasp predict grasp poses directly on point clouds.
ExampleA 640×480 depth image can be back-projected into more than 300,000 points at most. For 3D policies, the region outside the tabletop is usually cropped out first, and farthest point sampling then reduces the cloud to a few hundred or a few thousand points before it's fed into the encoder.
- Also called
- 3D Point Cloud, PCD
- Related
- Depth Map · Point Cloud Encoder · Farthest Point Sampling · Iterative Closest Point · PointNet / PointNet++ · 3D Diffusion Policy
- Sources
- Wikipedia: Point cloud
PCL 文档:The PCD (Point Cloud Data) file format (Chinese)
3D Diffusion Policy 项目页 (Chinese)