LLM-based Task Planning
大模型任务规划CommonUsing a large language or vision-language model to break a high-level instruction into a sequence of sub-tasks a robot can execute.
Task planning decides what to do first and what comes next. Traditionally this required hand-written formal descriptions in PDDL (Planning Domain Definition Language), which had to be rewritten for every new scenario. Starting around 2022, researchers began having large language models generate steps directly from an instruction: Google's SayCan has an LLM score candidate skills, then multiplies that score by a value function estimating how likely each skill is to succeed right now, to pick the next step; Code as Policies has the model write code that calls perception and control APIs; Inner Monologue feeds execution feedback back into the prompt for closed-loop replanning. LLMs readily produce steps that sound plausible but can't actually be executed, so LLM+P instead has the model translate only the natural-language goal into PDDL, which a classical planner then solves. Today, vision-language models (VLMs) are often used to plan directly from images — Gemini Robotics-ER, for instance — handing off each sub-task to a VLA model to execute.
ExampleGiven ‘I spilled my coke, can you bring me something to clean it up?,’ SayCan on a mobile manipulator selected, in sequence: find a sponge, pick up the sponge, bring it to you, done; the project page reported an 84% planning success rate and 74% execution success rate across 101 instructions.
- Also called
- LLM Planning, VLM Planning
- Related
- Task Planning · SayCan · Code as Policies · Planning Domain Definition Language · Gemini Robotics-ER · Dual-System Architecture (System 1 / System 2)
- Sources
- SayCan 项目主页 (Do As I Can, Not As I Say) (Chinese)
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (arXiv 2204.01691)
LLM+P: Empowering Large Language Models with Optimal Planning Proficiency (arXiv 2304.11477) - As of
- 2025-09