SayPlan
AdvancedA method that uses a 3D scene graph to let a large language model do long-horizon task planning across a large, multi-floor space.
SayPlan was released in July 2023 by researchers at the QUT Centre for Robotics, the University of Adelaide, and CSIRO's Data61, an oral-presentation paper at CoRL 2023. When using a large language model for robot task planning, writing every room and object in an entire building into the prompt quickly exceeds the context length. SayPlan represents the environment as a hierarchical 3D scene graph (floors, rooms, objects, and their states); the model is first shown only a collapsed, high-level version of this structure and performs a semantic search by 'expanding' and 'collapsing' nodes, keeping only the sub-graph relevant to the task. The actual path planning is left to classical algorithms like Dijkstra's, with the model only responsible for high-level steps. The generated plan is first checked in a scene-graph simulator, and infeasible steps — such as putting something into a cabinet that was never opened — are fed back to the model for revision. The test environments went up to 3 floors, 36 rooms, and 140 objects, with the final plan executed on a real mobile manipulation robot.
ExampleFor the instruction 'I'm hungry, bring some food to my desk,' SayPlan expands the kitchen in the scene graph to find the fridge and some food, generates a plan of 'walk to the fridge, open the fridge, take out an apple, walk to the desk, put it down,' and confirms in the simulator that every step is feasible.
- Also called
- SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
- Related
- 3D Scene Graph · LLM-based Task Planning · SayCan · Long-horizon Task · Replanning · Mobile Manipulation
- Sources
- SayPlan (arXiv 2307.06135)
SayPlan project page - As of
- 2023-09