Embodied AI Glossary中文

ALFRED

Advanced

A benchmark that has an agent complete multi-step household chores in a simulated home by following natural-language instructions.

ALFRED was proposed by Mohit Shridhar and colleagues (University of Washington, Carnegie Mellon, the Allen Institute for AI, NVIDIA), published at CVPR 2020, and built on the AI2-THOR simulator. It has 120 indoor scenes (30 each of kitchens, bathrooms, bedrooms, and living rooms), 8,055 expert demonstrations, and 25,743 English instructions, with each demonstration averaging about 50 steps; tasks fall into 7 categories, such as pick-and-place, stack-and-place, heat, cool, clean-and-place, and examining an object under a light. The agent sees only a first-person view and the instructions, and outputs discrete navigation and interaction actions, providing a pixel mask for the target object whenever it interacts. Metrics are task success rate and goal-condition success rate, plus versions weighted by path length. The original paper's baseline scored under 1% success on unseen scenes, versus about 91% for humans. Both ALFWorld and TEACh are built on top of it.

ExampleGiven the high-level goal “rinse off the mug and put it in the coffee maker,” paired with step-by-step instructions like “walk to the coffee maker on the right,” the agent must navigate, pick up the mug, rinse it at the sink, and then place it in the coffee maker, in sequence.

Also called
Action Learning From Realistic Environments and Directives
Related
Instruction Following · Long-horizon Task · AI2-THOR · ALFWorld · Household Tasks · Vision-and-Language Navigation
Sources
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks (arXiv 1912.01734)
ALFRED project site

See it in the full glossary →