Embodied AI Glossary中文

FoundationPose

Common

NVIDIA's general-purpose 6D pose model that estimates and tracks the pose of new objects without retraining.

FoundationPose is a 6D object pose (3D position plus 3D orientation) estimation and tracking model from Bowen Wen and colleagues at NVIDIA, published at CVPR 2024 as a Highlight paper. It targets objects the model has never seen during training, and needs no fine-tuning at test time: given either a CAD model or about 16 reference photos of the object, along with an RGB-D image and a detected region for the object, it outputs the pose. Its method scatters a large number of initial pose hypotheses uniformly around the object, compares a rendered image against the real one to refine each hypothesis, and then uses a ranking network to pick the best one; subsequent frames only need refinement, letting it track at roughly 32 Hz. Its training data is large-scale synthetic data generated with the help of large language models. As of March 2024, it ranked first on the BOP leaderboard for model-based pose estimation of novel objects, and an Isaac ROS version is also available.

ExampleAn arm needs to insert a workpiece it has never seen in training into a fixture: the workpiece's CAD model is scanned first, a segmentation model boxes the workpiece in the first frame, FoundationPose estimates its 6D pose and tracks it continuously, and the planner computes the grasp and insertion trajectory from that.

Also called
NVIDIA FoundationPose
Related
6D Object Pose Estimation · Pose Tracking · BOP (Benchmark for 6D Object Pose Estimation) · MegaPose · SAM-6D · NVIDIA
Sources
FoundationPose (arXiv:2312.08344)
NVlabs/FoundationPose (GitHub)
As of
2024-03

See it in the full glossary →