Embodied AI Glossary中文

Language-conditioned Policy

语言条件策略Common

A robot policy that takes a language instruction as input and produces different actions depending on what it says.

A policy is a mapping from observations to actions; a language-conditioned policy adds a natural-language instruction to that input, so a single network performs different tasks depending on the instruction, instead of training a separate model per task. Corey Lynch and Pierre Sermanet's language-conditioned imitation learning, proposed at Google in 2020 and published at RSS 2021, used one end-to-end network to learn pixel perception, language understanding, and continuous control together, drawing on large amounts of unlabeled “play” data, with less than 1% of it needing any language annotation. CALVIN (2021) is a benchmark built specifically to evaluate this kind of policy, requiring a robot to complete a long-horizon task made of a sequence of language instructions in order. Today's vision-language-action (VLA) models, such as RT-2 and π0, are essentially language-conditioned policies too, just built on a pretrained vision-language model backbone. The counterpart is a goal-conditioned policy, which specifies the task with a goal image instead of language.

ExampleThe same robot-arm policy pulls open a drawer when given the instruction “open the drawer,” and presses a button when given “press the green button” (tasks from the CALVIN benchmark).

Also called
Language-conditioned Imitation Learning
Related
Policy · Goal-conditioned Policy · Vision-Language-Action Model · Instruction Following · CALVIN Benchmark · Play Data
Sources
Language Conditioned Imitation Learning over Unstructured Data (Lynch & Sermanet)
CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

See it in the full glossary →