Embodied AI Glossary中文

Semantic Keypoints

语义关键点Advanced

Representing an object with a handful of meaningful points on it — a mug’s handle, a kettle’s spout — to make planning manipulation easier.

Semantic keypoints are a small number of 3D points on an object that carry clear meaning — for example, the center of a mug’s handle, the center of its base, or the heel of a shoe. Unlike 6D pose, this does not require an exact CAD template for every object; objects within the same category whose shapes vary a lot can still have corresponding points found on them, which makes the representation well suited to category-level generalization. kPAM, from Russ Tedrake’s group at MIT in 2019, applied this to robot manipulation: it first detects semantic 3D keypoints, then expresses the task as geometric constraints on those points (such as “hang the mug’s handle on the hook” or “press the mug’s base flat against the table”), and solves an optimization for the arm’s target pose, letting it manipulate new objects it has never seen. Recent work such as ReKep and OmniManip has vision-language models propose keypoints and constraints directly from an image, which is why the term “task keypoints” is also common.

ExampleIn kPAM, detecting only a few keypoints — the handle and the base — on mugs of many different, unseen shapes is enough to plan the motion for hanging each one on a mug rack.

Also called
Task Keypoints, Semantic 3D Keypoints
Related
Keypoint Detection · ReKep · OmniManip · 6D Object Pose Estimation · Category-Level Pose Estimation · Affordance
Sources
arXiv 1903.06684: kPAM: KeyPoint Affordances for Category-Level Robotic Manipulation

See it in the full glossary →