Pose Tracking
位姿跟踪AdvancedContinuously estimating an object’s 3D position and orientation across a video’s successive frames.
Pose means an object’s 3D position plus 3D orientation, 6 degrees of freedom in total, which is why it is also called 6D pose. 6D pose estimation usually computes the pose from scratch in a single frame; pose tracking instead uses the previous frame’s result and makes only a small correction in the new frame, which makes it faster and smoother, and suitable for real-time closed-loop control. NVIDIA’s FoundationPose (a CVPR 2024 Highlight paper) unifies estimation and tracking in one framework: given an RGB-D image, it needs only a CAD model or a handful of reference images for objects it has never seen, and refines the pose iteratively using a “render-and-compare” approach. Robots rely on it when grasping moving objects, adjusting an object in-hand, or doing visual servoing; if the object is occluded or moves too fast, tracking can be lost and a fresh global estimate is needed. Estimating a camera’s own pose in SLAM is also sometimes called tracking, but that is a different target.
ExampleWhile a robot unscrews a bottle cap, FoundationPose tracks the bottle’s 6D pose every frame, and the controller adjusts the gripper position accordingly.
- Also called
- 6D Pose Tracking, Object Pose Tracking
- Related
- 6D Object Pose Estimation · FoundationPose · Visual Servoing · In-hand Manipulation · Occlusion · Pose
- Sources
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects (arXiv 2312.08344)
FoundationPose 项目主页 (NVIDIA) (Chinese) - As of
- 2024-06