Embodied AI Glossary中文

Double-blind Pairwise Comparison

双盲成对比较Advanced

An evaluation method where a judge, without knowing which model is which, decides only which of two policies performed better.

Double-blind pairwise comparison is an evaluation method: two models or policies each run once under the same conditions, the evaluator doesn't know beforehand which is which, and simply judges which one did better; a large number of these “who won” records are then aggregated into a ranking. Chatbot Arena, in the large-language-model world, ranks models this way through anonymous head-to-head comparisons. A leading example in embodied AI is RoboArena (2025): 7 universities ran over 600 real-robot pairwise evaluations on the DROID platform to compare 7 generalist policies; evaluators could choose their own tasks and scenes, but always compared two policies blind. This avoids having to standardize scenes and success criteria in advance, and also reduces evaluators favoring their own institution's model; the win/loss records from many evaluation sites are then converted into scores using an Elo or Bradley-Terry model (a statistical model that estimates each competitor's ability score from win/loss outcomes). The paper argues that this kind of distributed evaluation produces more accurate and more scalable rankings than centralized evaluation.

ExampleAn evaluator gives the instruction “fold the towel” in their own lab, the system sends two anonymous policies, A and B, to perform it in turn, and the evaluator judges B to be better; that preference is recorded toward the leaderboard.

Also called
A/B Evaluation, Blind A/B Evaluation, Pairwise Preference Evaluation
Related
Elo Rating · RoboArena · Real-World Evaluation · Evaluation Protocol · DROID (Distributed Robot Interaction Dataset) · Benchmark
Sources
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)
As of
2025-06

See it in the full glossary →