RoboArena
CommonA generalist-robot-policy evaluation platform where multiple labs run blind head-to-head real-robot comparisons that get aggregated into a ranking.
RoboArena is a real-robot evaluation framework released in June 2025 by Karl Pertsch, Chelsea Finn, Sergey Levine, and colleagues from 7 academic institutions. It borrows the “arena” approach used to evaluate large language models: an evaluator freely sets up a scene and gives a task in their own lab, the system sends two anonymous policies to each run once from the same starting condition, and the evaluator provides a preference, a 0–100 progress score, and a written rationale; a Bradley-Terry model extended to account for task difficulty then aggregates a large number of these pairwise comparisons into a ranking. Policies connect as remote inference servers, so the evaluating site only needs a robot and a network connection. The first round standardized on the DROID platform (a Franka arm), and the paper reports that its ranking was closer to the reference ranking obtained from exhaustive testing than traditional single-lab, fixed-task evaluation was.
ExampleThe first evaluation round compared π0-flow-DROID, π0-FAST-DROID, and 5 other DROID policies built on PaliGemma with different action representations, completing over 600 blind pairwise comparisons across the 7 participating institutions.
- Also called
- Robo Arena, RoboArena Distributed Real-Robot Evaluation
- Related
- Double-blind Pairwise Comparison · Elo Rating · Real-World Evaluation · DROID (Distributed Robot Interaction Dataset) · Progress Score · Policy Server (Remote Inference)
- Sources
- RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)
RoboArena project page - As of
- 2025-11