Embodied AI Glossary中文

Transparent & Reflective Object Perception

透明/反光物体感知Advanced

Getting a robot to correctly see glass, metal, and other objects that ordinary depth cameras get wrong.

This refers to detecting, segmenting, and estimating the 3D shape of objects such as glass, clear plastic, and mirror-like metal. Common depth cameras measure range using reflected infrared light, but that light passes straight through transparent objects or bounces away off a mirror-like surface, so the resulting depth map ends up with holes (pixels with no reading) or, worse, reports the distance to whatever is behind the object, causing a robot that grasps based on that depth to miss entirely. A typical approach uses a neural network to predict a mask of the transparent region, surface normals, and occlusion boundaries from the color image, and then performs depth completion (filling in the missing depth); ClearGrasp, from Sajjan, Andy Zeng, Shuran Song, and colleagues in 2019, is representative work in this area. Glasses and bottles are everywhere in household, lab, and retail settings, so this is a problem any real grasping system has to confront.

ExampleClearGrasp estimates surface normals, a mask, and occlusion boundaries for a transparent object from a single RGB-D image, corrects the depth, and hands it to a grasping algorithm, improving a robot arm’s success rate at grasping transparent objects.

Also called
Transparent Object Depth Estimation, Highly Reflective Object Perception, Transparent Object Perception
Related
Depth Completion · Depth Holes · Depth Camera · Surface Normal Estimation · Grasp Pose Detection · Active Stereo
Sources
ClearGrasp: 3D Shape Estimation of Transparent Objects for Manipulation (arXiv:1910.02550)

See it in the full glossary →