ALFWorld
AdvancedA text-game version of the ALFRED household tasks, letting an agent learn in text first and then act on the actual visuals.
ALFWorld was proposed by Shridhar, Xingdi Yuan, and colleagues (University of Washington, Microsoft Research, Carnegie Mellon), published at ICLR 2021. It describes each ALFRED scene in PDDL (Planning Domain Definition Language) and uses Microsoft's TextWorld engine to generate an equivalent text-adventure game: the agent reads a text description of which furniture and items are in the room and acts through text commands like go to desk 1. Tasks fall into 6 categories, with 3,553 games used for training. The authors' BUTLER agent first learns a high-level policy in the text environment, then adds visual recognition (Mask R-CNN) and a low-level control module to transfer it to execution on AI2-THOR's actual visuals, training about 7 times faster than learning directly from visuals alone while also generalizing better. After large language models took off, agent frameworks such as ReAct adopted it as a standard test environment.
ExampleFor the task “examine the alarm clock under the desk lamp,” the text version is solved by entering go to desk 1, take alarmclock 2 from desk 1, and use desklamp 1 in sequence.
- Also called
- Aligning Text and Embodied Environments for Interactive Learning
- Related
- ALFRED · AI2-THOR · LLM-based Task Planning · Planning Domain Definition Language · Long-horizon Task · EmbodiedBench
- Sources
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning (arXiv 2010.03768)
ALFWorld project site
ReAct: Synergizing Reasoning and Acting in Language Models (arXiv 2210.03629)