Embodied AI Glossary中文

Keypoint Detection

关键点检测Common

Locating a handful of pre-defined, meaningful points in an image, like a wrist, a cup's handle, or a box's corner.

Keypoint detection locates a small number of points with predefined meaning in an image, such as a wrist, a cup's handle, or a box's corner. It's more precise than a bounding box and lighter-weight than a segmentation mask; it also differs from feature points like SIFT or ORB, which only need to be locally recognizable and carry no fixed meaning. The most common case is human body keypoints: COCO labels each point as ‘not labeled,’ ‘occluded,’ or ‘visible,’ and scores predictions with OKS similarity. In robot manipulation, MIT's kPAM (2019) represents a whole category of objects with a few 3D semantic keypoints, so swapping in a differently shaped cup still works under the same rule — handling shape variation within a category better than estimating a single 6D pose would.

ExampleTo hang various mugs on a mug rack, the robot first detects each mug's base, rim, and handle as 3D keypoints, then plans motion so the handle lines up with the hook — the same rule works even though the mugs differ in size and shape.

Also called
Keypoints, Keypoint Localization
Related
Human Pose Estimation · Semantic Keypoints · Feature Points · 6D Object Pose Estimation · ReKep · Tracking Any Point
Sources
COCO Data Format(keypoints 标注格式) (Chinese)
kPAM: KeyPoint Affordances for Category-Level Robotic Manipulation (arXiv:1903.06684)

See it in the full glossary →