ERQA
ERQA 具身推理问答基准AdvancedA 400-question multiple-choice benchmark from Google DeepMind, pairing images and text to test embodied reasoning.
ERQA is a benchmark Google DeepMind open-sourced in March 2025 alongside Gemini Robotics, used to measure a multimodal model's embodied reasoning: its ability to understand real physical scenes and make judgments for robot action. It has 400 multiple-choice questions (options A–D) that mix images and text, with scenes mostly drawn from real robot-relevant environments; question types include spatial reasoning, trajectory reasoning, action reasoning, state estimation, pointing, multi-view reasoning, and task reasoning, and 28% of questions require looking at multiple images. Because it's purely multiple-choice, it needs no robot or simulator, and any vision-language model can run it directly. In the Gemini Robotics technical report, Gemini 2.0 Pro Experimental scored 48.3%, GPT-4o scored 47.0%, and Claude 3.5 Sonnet scored 35.5%, showing that models at the time were still far from reliable embodied reasoning. Evaluation platforms such as Embodied Arena have also incorporated it.
ExampleA typical question shows a photo of a robot arm holding a cup of water and asks, “to pour the water into the bowl next to it, which direction should the gripper rotate next?” with four options, A through D, to choose from.
- Also called
- Embodied Reasoning Question Answer Benchmark
- Related
- Embodied Reasoning · Gemini Robotics-ER · Visual Question Answering · Multimodal Large Language Model · Embodied Arena · VSI-Bench
- Sources
- embodiedreasoning/ERQA GitHub
Gemini Robotics: Bringing AI into the Physical World (arXiv 2503.20020) - As of
- 2025-03