Embodied AI Glossary中文

Category-Level Pose Estimation

类别级位姿估计Advanced

Estimates 6D pose and size for an unseen object in a known category, without needing that exact object’s CAD model.

Instance-level 6D pose estimation requires an accurate 3D model of each specific object beforehand, so it can only handle the exact objects seen during training. Category-level pose estimation relaxes this to: knowing only which category an object belongs to — cup, bowl, laptop, and so on — and estimating position, orientation, and 3D size for a new, unseen instance of that category. The difficulty is that objects in the same category can vary a lot in shape. He Wang and colleagues in Leonidas Guibas’s group at Stanford introduced NOCS (Normalized Object Coordinate Space) at CVPR 2019, which has a network map every pixel to a standardized coordinate shared across the category, then combines this with a depth map to solve for pose and size; they also released the CAMERA (synthetic) and REAL275 (real) datasets that are widely used since. This suits household tasks like “pick up any mug”; zero-shot methods like FoundationPose take a different approach, requiring a model or reference image of the new object instead.

ExampleA robot sees a mug on the table it has never seen before; a network maps the mug’s pixels to the “mug” category’s standard coordinates, aligns this with the depth point cloud, and computes the mug’s position, orientation, and size, from which it plans a pose for grasping the handle.

Also called
Category-Level 6D Pose, Category-Level 6D Object Pose and Size Estimation
Related
6D Object Pose Estimation · Normalized Object Coordinate Space · FoundationPose · Pose Tracking · Object Generalization · BOP (Benchmark for 6D Object Pose Estimation)
Sources
Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation (arXiv 1901.02970, CVPR 2019)

See it in the full glossary →