Contact-GraspNet
AdvancedNVIDIA’s grasping network that generates 6-DoF grasps directly from a depth point cloud in cluttered scenes.
Contact-GraspNet was proposed by Sundermeyer, Mousavian, Triebel, and Fox at NVIDIA, published at ICRA 2021. It targets two-finger parallel-jaw grippers: given a depth image plus camera intrinsics (or a point cloud directly), with an optional object segmentation mask, it outputs a batch of 6-DoF grasp poses with confidence scores end to end. Its key design is treating every observed 3D point as a possible fingertip contact point, which leaves only the gripper’s 3D orientation and opening width — 4 degrees of freedom — left to predict, greatly easing the learning problem. The model doesn’t distinguish between object categories, and it was trained on 17 million simulated grasps (grasp annotations from the ACRONYM dataset); the paper reports over 90% success grasping unseen objects in cluttered real-robot scenes. It’s commonly used as the grasping module right after segmentation or open-vocabulary detection.
ExampleA table is cluttered with cups, boxes, and toys; a segmentation model first cuts out the target cup, and the depth image plus the cup’s mask are fed into Contact-GraspNet, which picks the highest-confidence grasp among the candidates that land on the cup for motion planning to execute.
- Also called
- Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes
- Related
- Grasp Pose Detection · AnyGrasp · ACRONYM: A Large-Scale Grasp Dataset Based on Simulation · GraspNet-1Billion · Point Cloud · Parallel Jaw Gripper
- Sources
- Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes (arXiv 2103.14127)
NVlabs/contact_graspnet (GitHub) - As of
- 2021-03