3D Vision
3D视觉CommonThe umbrella term for techniques that recover an object's or scene's 3D geometry from images or sensor data.
3D vision is the branch of computer vision that studies 3D geometry, with the goal of recovering depth, point clouds, meshes, and object poses — in short, where things are and what shape they have. It splits into two families by how the data is captured: active methods emit their own signal to measure distance, such as structured light, time-of-flight cameras, and lidar; passive methods use only an ordinary camera, relying on stereo disparity, multi-view structure from motion, or neural-network monocular depth estimation. Common tasks include depth estimation, 3D reconstruction, point cloud segmentation and registration, 6D pose estimation, and 3D object detection. A robot reaching for, avoiding, or placing an object needs precise distance information that a plain 2D image can't provide, which is why grasp planning, 3D Diffusion Policy, and 3D VLA models all depend on 3D vision; feed-forward reconstruction models such as DUSt3R and VGGT have made it substantially easier to use.
ExampleBefore an arm grasps a cup, a depth camera turns the scene into a point cloud, an algorithm segments out the cup within that point cloud and estimates the position and orientation of its handle, and that result is converted into the 3D coordinates the gripper needs to reach.
- Also called
- 3D Perception
- Related
- Point Cloud · Depth Camera · Stereo Camera · 6D Object Pose Estimation · Feed-Forward 3D Reconstruction · 3D VLA
- Sources
- Wikipedia: 3D reconstruction
Wikipedia: Computer stereo vision