CoPa
AdvancedA training-free framework where GPT-4V finds relevant object parts and writes spatial constraints to follow open-ended manipulation instructions.
CoPa was proposed in March 2024 by Yang Gao's group at Tsinghua University, the Shanghai Qi Zhi Institute, Shanghai Jiao Tong University, and the Shanghai AI Lab. Rather than training a robot policy, it has a foundation model (a general-purpose large model pretrained on massive data) directly supply the geometric information manipulation needs. A manipulation is split into two steps. First, task-oriented grasping: GPT-4V, combined with Set-of-Mark visual prompting, picks where to grasp by narrowing from the whole object down to a specific part, and GraspNet then generates the grasp pose. Second, task-oriented motion planning: a vision-language model identifies the task-relevant parts (such as a hammer's head and a nail), writes out the spatial constraints they should satisfy, solves for the target pose after grasping, and hands it to a motion planner to execute. Like VoxPoser and ReKep, it belongs to the “large model supplies constraints, classical planner executes” approach, and it can handle open-ended instructions and unseen objects.
ExampleGiven the instruction “hammer in the nail,” CoPa first has GPT-4V select the hammer's handle as the grasp part, then locates the hammer's striking face and the nail, constrains the striking face to align with the nail with the swing direction matching the nail's axis, solves for the hammer's pose right before striking, and executes it.
- Also called
- Spatial Constraints of Parts, CoPa: General Robotic Manipulation through Spatial Constraints of Parts with Foundation Models
- Related
- ReKep · VoxPoser · OmniManip · Visual Prompting (Set-of-Mark) · Task-Oriented Grasping · Affordance
- Sources
- CoPa (arXiv:2403.08248)
CoPa 项目主页 (Chinese) - As of
- 2024-03