Embodied AI Glossary中文

Projection / Back-Projection

投影与反投影Common

Projection maps a 3D point onto a pixel; back-projection uses a pixel plus its depth to recover the 3D point.

Projection and back-projection are the two directions of the pinhole camera model. Projection: given a 3D point (X, Y, Z) in the camera's coordinate frame and the intrinsics, compute which pixel (u, v) it lands on. Back-projection: a single pixel only defines a ray, so recovering the 3D point also needs that pixel's depth Z: X = (u − cx)·Z/fx, Y = (v − cy)·Z/fy. Doing this for every pixel of a depth map produces a point cloud, and both Open3D and the RealSense SDK provide ready-made functions for it. A few practical points: raw depth values are usually stored as integers (RealSense D400 cameras default to 1 unit = 1 millimeter) and must be converted to meters; color and depth images from different sensors need to be aligned first; and the result must then be multiplied by the camera's extrinsics to land in the robot's base frame.

ExampleA VLM points at a cup's handle at pixel (412, 305) on the color image. Looking up the aligned depth gives 0.52 meters; back-projection with the intrinsics gives the 3D point in the camera frame, and the hand-eye-calibrated extrinsics then convert it into the arm's base frame as the grasp target.

Also called
Depth-to-Point-Cloud Conversion, Unprojection, Deprojection
Related
Pinhole Camera Model · Camera Intrinsics · Depth Map · Point Cloud · Depth-to-Color Alignment · Camera Extrinsics
Sources
Intel RealSense Wiki: Projection in RealSense SDK 2.0
Open3D API: open3d.geometry.PointCloud (create_from_depth_image)

See it in the full glossary →