Embodied Arena
Embodied Arena 具身评测竞技场AdvancedAn evaluation platform that plugs in benchmarks for embodied question answering, navigation, and task planning under one live leaderboard.
Embodied Arena is an embodied-AI evaluation platform released jointly by over a dozen institutions in China and abroad, including Tianjin University and Huawei's Noah's Ark Lab, with its paper posted to arXiv in September 2025. It addresses the problem that embodied AI has many benchmarks but each operates in isolation, making models hard to compare directly and leaving unclear what capabilities embodied AI actually requires. The platform first defines a capability taxonomy with three layers — perception, reasoning, and task execution — broken into 7 core capabilities and 25 finer-grained dimensions. It then plugs 22 existing benchmarks into a unified evaluation backend, covering 2D/3D embodied question answering (such as OpenEQA, VSI-Bench, and ERQA), navigation (such as R2R-CE and HM3D), and task planning (such as EB-ALFRED and EB-Habitat), evaluating over 30 models from more than 20 institutions; it also uses an LLM-driven pipeline to automatically generate new evaluation data, keeping the question pool continuously updated. Results are published as three live leaderboards, viewable either by benchmark or by capability dimension.
ExampleTo check a multimodal model's spatial-reasoning ability, a user can simply look at that dimension's aggregate score in Embodied Arena's capability view, instead of running VSI-Bench, ERQA, and other benchmarks separately.
- Also called
- A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- Related
- Benchmark · EmbodiedBench · ERQA · OpenEQA (Open-Vocabulary Embodied Question Answering Benchmark) · VSI-Bench · Embodied Question Answering
- Sources
- Embodied Arena (arXiv 2509.15273)
Embodied Arena 论文 HTML 版(作者单位与基准列表) (Chinese) - As of
- 2025-09