Embodied AI Glossary中文

Benchmark Saturation

基准饱和Advanced

When leading models' scores on a benchmark all cluster near the maximum, so it can no longer distinguish good methods from bad ones.

A benchmark — a shared test with fixed tasks and an evaluation protocol — tends to saturate the longer it's used: scores from different groups converge toward the ceiling, with gaps shrinking to a percentage point or two, sometimes within the range of random noise. This can happen because methods genuinely improved, but it can equally happen because everyone has repeatedly tuned against the same test set, or because the test scenes are too similar to the training data, letting a model score well through memorization. The Dynabench paper (2021) in NLP made exactly this point: models quickly achieve excellent benchmark scores yet fail on simple adversarial examples. A typical case in embodied AI is LIBERO, where multiple VLA models now score above 90% success under the standard setup. Once a benchmark saturates, the community usually releases a harder or perturbed successor, or shifts to real-robot evaluation and generalization/robustness evaluation. A small lead on a saturated benchmark carries limited weight.

ExampleLIBERO-PRO (2025) shows that a model scoring above 90% success on standard LIBERO drops to 0.0% success once objects are swapped, initial states changed, instructions reworded, or the environment changed — indicating that the high score largely reflected memorization of training trajectories and scene layouts.

Also called
Leaderboard Saturation
Related
Benchmark · LIBERO Benchmark · LIBERO-PRO · LIBERO-Plus · Leaderboard Chasing · Generalization / Robustness Evaluation
Sources
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization (arXiv 2510.03827)
Dynabench: Rethinking Benchmarking in NLP (arXiv 2104.14337)
Stanford HAI: The 2025 AI Index Report
As of
2025-10

See it in the full glossary →