Embodied AI Glossary中文

RoboArena

Common

A generalist-robot-policy evaluation platform where multiple labs run blind head-to-head real-robot comparisons that get aggregated into a ranking.

RoboArena is a real-robot evaluation framework released in June 2025 by Karl Pertsch, Chelsea Finn, Sergey Levine, and colleagues from 7 academic institutions. It borrows the “arena” approach used to evaluate large language models: an evaluator freely sets up a scene and gives a task in their own lab, the system sends two anonymous policies to each run once from the same starting condition, and the evaluator provides a preference, a 0–100 progress score, and a written rationale; a Bradley-Terry model extended to account for task difficulty then aggregates a large number of these pairwise comparisons into a ranking. Policies connect as remote inference servers, so the evaluating site only needs a robot and a network connection. The first round standardized on the DROID platform (a Franka arm), and the paper reports that its ranking was closer to the reference ranking obtained from exhaustive testing than traditional single-lab, fixed-task evaluation was.

ExampleThe first evaluation round compared π0-flow-DROID, π0-FAST-DROID, and 5 other DROID policies built on PaliGemma with different action representations, completing over 600 blind pairwise comparisons across the 7 participating institutions.

Also called
Robo Arena, RoboArena Distributed Real-Robot Evaluation
Related
Double-blind Pairwise Comparison · Elo Rating · Real-World Evaluation · DROID (Distributed Robot Interaction Dataset) · Progress Score · Policy Server (Remote Inference)
Sources
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)
RoboArena project page
As of
2025-11

See it in the full glossary →