Embodied AI Glossary中文

Evaluation Protocol

评测协议Advanced

The set of rules specifying under what conditions a policy is tested, how many trials, and what counts as success.

An evaluation protocol is a set of rules that spells out exactly how a policy is tested: which scenes and objects it's tested on, how the initial state is arranged, how many trials each condition gets, the maximum duration of a trial, what counts as success, whether human intervention is allowed mid-trial, and what metrics summarize the results. Robot evaluation results are very sensitive to these details — the same policy can get a very different success rate depending on object placement or how loosely success is defined — and if a paper doesn't spell out its protocol, readers can't tell whether the results are reproducible or comparable across methods. A 2024 paper by Kress-Gazit and colleagues, “Robot Learning as an Empirical Science,” recommends clearly reporting experimental conditions and success criteria, supplementing success rate with other metrics, doing statistical analysis, and qualitatively describing failure modes. Simulation benchmarks (such as LIBERO and CALVIN) usually come with a fixed protocol built in; real-robot evaluation more often relies on practices like a fixed table of initial positions, alternating between methods during testing, and double-blind pairwise comparison to keep things fair.

ExampleA protocol might read: test each task 20 times, drawing the object's initial position in turn from 20 pre-marked points, cap each trial at 60 seconds, count success only when the object is fully inside the box and the gripper has released, and allow no manual resets mid-trial.

Also called
Testing Protocol
Related
Benchmark · Real-World Evaluation · Simulation-Based Evaluation · Success Rate · Statistical Rigor in Policy Evaluation (Confidence Intervals / Sequential Testing / Multiple Seeds) · Double-blind Pairwise Comparison
Sources
Robot Learning as an Empirical Science: Best Practices for Policy Evaluation (arXiv 2409.09491)
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)

See it in the full glossary →