Embodied AI Glossary中文

Dense Correspondence

稠密对应Advanced

Finds, for every pixel in one image, the matching location in another image.

Dense correspondence means finding a match for every pixel (or nearly every pixel) between two images, as opposed to sparse matching, which matches only a small number of feature points. It comes in two flavors: geometric correspondence finds where the same physical point appears from a different viewpoint or at a different time — optical flow between neighboring video frames is one example; semantic correspondence finds parts with the same meaning across different objects, like the handles of two differently shaped cups. Recent work often uses nearest-neighbor matching on features from pretrained vision models directly, such as DINO-family features, or the features DIFT extracts from a diffusion model at NeurIPS 2023, with no dedicated training needed. In robotics, this lets a grasp point or keypoint from a demonstration transfer to a new object; MIT’s 2018 Dense Object Nets used self-supervised dense descriptors to transfer grasps between objects of the same category.

ExampleIn a demonstration, a person grasps the handle of a red cup; swapping in a never-before-seen blue mug, dense correspondence using DINOv2 features finds the matching handle pixels on the blue mug, and combined with depth this gives the grasp location.

Also called
Semantic Correspondence, Dense Matching
Related
Feature Matching · Semantic Keypoints · Optical Flow · DINOv2 · Keypoint Detection · Neural Descriptor Fields
Sources
Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation (arXiv 1806.08756)
Emergent Correspondence from Image Diffusion (DIFT, arXiv 2306.03881)

See it in the full glossary →