Embodied AI Glossary中文

Inner Monologue

Advanced

Feeds environment feedback back into a large language model as text, letting the robot adjust its plan as it goes.

Inner Monologue was released in July 2022 by Wenlong Huang, Fei Xia, Brian Ichter, and colleagues at Robotics at Google, published at CoRL 2022. At the time, work such as SayCan already used large language models to break a high-level instruction into a sequence of skills, but the plan was fixed once made, and the model had no way of knowing when something went wrong during execution. Inner Monologue needs no extra training: it writes several kinds of feedback back into the LLM's prompt in natural language — whether a skill succeeded (success detection), what's in the scene (passive or active scene description), and any additional human instructions — forming an “inner monologue” that the LLM uses to decide its next step, retry, or revise the plan. Across three settings — simulated and real tabletop object rearrangement, and long-horizon mobile manipulation in a real kitchen — this closed-loop language feedback clearly raised the instruction-completion rate. It is an early representative example of using an LLM for closed-loop robot planning.

ExampleFor example, when the robot's grasp fails, a success detector writes back “action failed”; reading this, the LLM schedules a re-grasp attempt instead of just moving on to the next step.

Also called
Embodied Reasoning through Planning with Language Models
Related
SayCan · LLM-based Task Planning · Success Detector · Closed-loop Control · Long-horizon Task · Code as Policies
Sources
Inner Monologue (arXiv 2207.05608)
Inner Monologue project page
Inner Monologue (PMLR v205, CoRL 2022)
As of
2022-12

See it in the full glossary →