Embodied AI Glossary中文

RobotArena Infinity (RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation)

RobotArena ∞Advanced

A benchmark that auto-converts real robot videos into simulated scenes, then ranks VLA models by scoring and human voting.

RobotArena ∞ is a robot-policy evaluation framework released in October 2025 by Katerina Fragkiadaki's group at Carnegie Mellon University. Real-robot evaluation is labor-intensive, slow, unsafe, and hard to reproduce, so this framework instead takes real manipulation videos from public datasets such as BridgeData V2, DROID, and RH20T and, using object segmentation, single-image-to-3D generation, and background inpainting models together with camera calibration and system identification, automatically converts them into digital-twin scenes inside the Genesis simulator, where different VLA models can then execute. Scoring runs two ways: a vision-language model gives a per-frame task-progress score, and crowd reviewers watch two execution videos double-blind and vote for the better one, with results aggregated into an Elo-style ranking using a Bradley-Terry model. It also systematically swaps backgrounds, changes colors, and moves objects to test robustness. The paper evaluated 6 policies — Octo, CogACT, π0, and X-VLA among them — and collected more than 8,500 human preference comparisons.

ExampleConvert a real tabletop-manipulation video from DROID into a Genesis scene, let π0 and Octo each execute the same instruction inside it, have crowd reviewers watch both replays double-blind and vote for the better one to feed the leaderboard, then swap the background and rerun to see if the ranking changes.

Also called
RobotArena Infinity, RobotArena ∞
Related
Simulation-Based Evaluation · Real-to-Sim · Double-blind Pairwise Comparison · Elo Rating · Digital Twin · RoboArena
Sources
arXiv 2510.23571 - RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation
RobotArena ∞ 论文 HTML 版 (Chinese)
As of
2026-03

See it in the full glossary →