Embodied AI Glossary中文

Embodied Question Answering

具身问答EQAAdvanced

A task where an agent moves through a 3D environment on its own to gather information, then answers a question.

Embodied question answering was proposed by Abhishek Das, Dhruv Batra, and colleagues in late 2017 (an oral presentation at CVPR 2018): an agent is placed at a random position in a 3D house environment and given a question, such as “what color is the car,” and it must navigate on its own from a first-person view, find the relevant object, and then answer. The first dataset, EQA v1, was built on the House3D simulator, with questions covering location, color, room color, and prepositional relations. This task combines active perception, language understanding, object-goal navigation, commonsense reasoning, and language grounding all at once, which is why it is often used as a comprehensive benchmark for embodied AI. Later variants added multiple targets and real scanned environments; Meta FAIR's OpenEQA, released at CVPR 2024, extended it to an open vocabulary and split it into two settings, answering from memory and answering after active exploration. In the large-model era, it is commonly tackled with a vision-language model combined with frontier exploration.

ExampleAsked “how many chairs are in the kitchen,” the agent needs to start from the living room, find the kitchen, count the chairs, and then give the answer.

Also called
EQA, Embodied QA
Related
Embodied Interaction · Visual Question Answering · OpenEQA (Open-Vocabulary Embodied Question Answering Benchmark) · Active Exploration · Vision-and-Language Navigation · Embodied Memory
Sources
Embodied Question Answering (Das et al., CVPR 2018)
EmbodiedQA project page
OpenEQA: Embodied Question Answering in the Era of Foundation Models
As of
2024-06

See it in the full glossary →