Real-World Evaluation
真机评测EssentialRunning a policy on an actual robot repeatedly and recording success rate and other outcomes.
Real-world evaluation means running a policy on an actual robot in an actual scene, recording whether each attempt succeeded, how far it got, and whether a human had to step in. No matter how good a simulation is, a sim-to-real gap remains, so this is the ultimate test of an embodied model. The difficulty is cost and reproducibility: a person has to place objects and reset the scene every time, and results are judged by hand; a slight difference in lighting, placement, or the robot's own state can shift the score, making it hard to compare numbers across different labs. For this reason, papers need to spell out their evaluation protocol clearly — the tasks, number of trials, initial conditions, and success criteria. To make results more trustworthy and more scalable, approaches such as RoboArena's distributed, double-blind pairwise comparisons have emerged, alongside using simulated evaluation (like SimplerEnv) or a world model to predict real-robot performance instead.
ExampleRoboArena ran more than 600 double-blind, pairwise real-robot comparisons of 7 general-purpose policies across DROID robot setups at 7 universities, then aggregated the results into a ranking.
- Also called
- real-robot evaluation, real-robot testing
- Related
- Simulation-Based Evaluation · Evaluation Protocol · Success Rate · Double-blind Pairwise Comparison · RoboArena · Sim-to-Real Gap (Reality Gap)
- Sources
- RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)
Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941) - As of
- 2025-06