Affordance Detection
可供性检测AdvancedFinds where on an object you can grip, press, or pour from an image or point cloud, and marks it to a specific region.
Affordance refers to the actions an object offers an agent — a handle affords gripping, a button affords pressing. Affordance detection pins this possibility down to specific pixels or points: given an RGB image, depth image, or point cloud, it outputs an action label or heatmap (an affordance map) for each region. An early landmark, AffordanceNet (Do et al., ICRA 2018), detects objects and simultaneously assigns each pixel of the object its most likely affordance label. “Affordance grounding” puts more emphasis on finding a region for a given action word — for instance, Luo and colleagues’ AGD20K dataset (CVPR 2022, over 20,000 images across 36 affordance classes) learns from third-person human-object interaction images and transfers this knowledge to object-only images. This fills the gap between “recognizing an object” and “knowing how to operate it,” and the result is often used as an intermediate representation for grasp planning or a policy. Recent work also uses vision-language models to predict actionable points directly, such as RoboPoint.
ExampleGiven the instruction “pour a cup of water,” affordance detection on an image of a kettle labels the handle a “grip” region and the spout a “pour” region, so the grasping module only samples grasp poses on the handle.
- Also called
- Affordance Grounding, Affordance Map, Affordance Prediction
- Related
- Affordance · Grasp Pose Detection · Task-Oriented Grasping · RoboPoint · VRB · Intermediate Representation
- Sources
- AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection (arXiv:1709.07326)
Learning Affordance Grounding from Exocentric Images (CVPR 2022, arXiv:2203.09905)