Embodied AI Glossary中文

Instruction Following

指令跟随Common

A robot understanding a natural-language instruction from a person and actually carrying it out in the real world.

In embodied AI, instruction following means a robot carries out the corresponding action or task based on a natural-language instruction, sometimes paired with an image or gesture, rather than only executing a fixed, pre-written program. Early research mostly used short instructions like “push the red block to the left”; the field has since moved toward open-ended instructions — Physical Intelligence's Hi Robot (2025), for example, has to handle a complex request like “make me a vegetarian sandwich,” as well as real-time corrections mid-execution such as “that's not trash.” It uses a hierarchical structure: a high-level vision-language model understands the instruction and any feedback and decides what to do next, while a low-level policy executes the specific actions. Instruction following determines whether an ordinary person can assign a robot a task without writing code, and it is a main dimension for evaluating a VLA (vision-language-action) model's semantic generalization.

ExampleA user says, “throw away the trash on the table, but don't touch that cup.” The robot has to identify which items are trash, avoid the cup, and pick the trash up piece by piece to put it in the bin.

Related
Language-conditioned Policy · Vision-Language-Action Model · Hi Robot · Language Grounding · Semantic Generalization · Language Corrections
Sources
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
As of
2025-02

See it in the full glossary →