VoxPoser
CommonA zero-shot method where a large model writes code to build a 3D value map, which a planner turns into a trajectory.
VoxPoser was released in July 2023 by Fei-Fei Li and Jiajun Wu's groups at Stanford, working with Yunzhu Li at UIUC, and received an oral presentation at CoRL 2023, with Wenlong Huang as first author. Large language models know a lot of commonsense, but not where in 3D space an action should land. VoxPoser has an LLM write code, from an instruction, that calls a vision-language model to locate relevant objects, then assembles an “affordance map” (where to go) and a “constraint map” (where to avoid) on a voxel grid (space cut into small cubes); these maps become cost functions handed to a motion planner, which computes the end-effector's trajectory. The whole process needs no task-specific policy training, so it's zero-shot, and it replans in closed loop during execution, handling objects that get moved. It's a landmark example of the “large model plus intermediate representation plus classical planning” approach, and later work like ReKep continues in the same direction.
ExampleGiven the instruction “open the top drawer, careful of the vase nearby,” the LLM-written code first locates the top drawer's handle and marks the area around it with high value; it then locates the vase and marks the nearby region as high cost; the planner uses this to compute a trajectory to the handle that avoids the vase.
- Also called
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- Related
- LLM-based Task Planning · Code as Policies · Affordance · Intermediate Representation · ReKep · Motion Planning
- Sources
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models (arXiv 2307.05973)
VoxPoser 项目页 (Chinese) - As of
- 2023-11