Embodied AI Glossary中文

Benchmark

基准测试Essential

A fixed set of tasks, data, and scoring rules that let different methods be compared under the same conditions.

The term “benchmark” originates in computing, referring to a standardized set of tests used to measure the relative performance of something. In embodied AI, a benchmark typically bundles a fixed set of tasks, a simulator or real-world setup, demonstration data (if any), an evaluation protocol, and a metric — usually success rate. Common simulated benchmarks include LIBERO, CALVIN, SimplerEnv, and RoboTwin, alongside real-robot evaluation networks such as RoboArena. Its value is reproducibility and the ability to compare methods head to head; the risk is that methods can be tuned specifically to score well on the leaderboard (“benchmark hacking”) without that reflecting real-world usefulness, and once leading methods get close to a perfect score, the benchmark stops being able to distinguish between them — what's called benchmark saturation.

ExampleVLA papers commonly report average success rate across LIBERO's four task suites, placing their numbers in the same table as baselines such as OpenVLA for comparison.

Also called
leaderboard
Related
Baseline · Evaluation Protocol · Success Rate · LIBERO Benchmark · Benchmark Saturation · Leaderboard Chasing
Sources
Wikipedia: Benchmark (computing)
LIBERO (NeurIPS 2023 Datasets and Benchmarks Track)

See it in the full glossary →