Elo Rating
Elo 评分CommonA scoring method that estimates each player's or policy's relative strength from a series of head-to-head wins and losses.
Elo rating was designed by physics professor and chess master Arpad Elo to rank chess players, and the US Chess Federation adopted it in 1960. Each player gets a single number; the gap between two ratings predicts the expected win probability, and after each game a rating is nudged by (actual result minus expected result) times a constant K. Mathematically, Elo is a special case of the Bradley-Terry model, which writes the probability that A beats B as a sigmoid function of the difference in their abilities. Embodied AI borrows this to solve a real problem: success rates reported by different labs on different tasks cannot be compared directly. Instead, two policies are compared head-to-head on the same task, and aggregating many such comparisons produces a leaderboard. RoboArena found that plain Elo rankings get distorted when tasks vary widely in difficulty, so it uses a Bradley-Terry variant that adds a task-difficulty parameter.
ExampleRoboArena has evaluators at different sites blindly compare two policies on tasks of their own choosing and judge which one did better, then aggregates over 600 real-robot comparisons into a leaderboard of seven general-purpose policies using a modified Bradley-Terry model.
- Also called
- Bradley-Terry Model, BT Model, Elo Rating System
- Related
- Double-blind Pairwise Comparison · RoboArena · Real-World Evaluation · Success Rate · Benchmark
- Sources
- RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)
Elo rating system - Wikipedia - As of
- 2025-11