Embodied AI Glossary中文

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

VLA-ArenaAdvanced

Peking University's open-source VLA evaluation framework, grading capability boundaries across task, language, and vision axes.

VLA-Arena is an open-source VLA (vision-language-action model) evaluation framework released by Peking University's PKU-Alignment team, posted to arXiv in December 2025 and later accepted at ICML 2026, with simulation built on LIBERO and robosuite. It breaks difficulty into three independent axes: task structure, language instructions, and visual observations. There are 11 task suites totaling 170 tasks, split into four categories — safety, distractors, extrapolation, and long-horizon — with each suite graded L0–L2, where fine-tuning is only allowed on L0; language (W0–W4) and vision (V0–V4) perturbations can then be layered onto any task. Using this setup, the authors find that current VLAs tend to memorize their training tasks, understand vision only shallowly, and often disregard safety constraints.

ExampleThe paper found that a model ranking near the top on L0 could be overtaken by other models once moved to L1 or L2; the authors treat this kind of rank reversal as evidence that the three difficulty levels each carry independent information.

Also called
VLA-Arena Benchmark
Related
Vision-Language-Action Model · LIBERO Benchmark · robosuite · Generalization / Robustness Evaluation · Embodied Safety · Benchmark
Sources
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models (arXiv 2512.22539)
VLA-Arena 项目主页 (Chinese)
As of
2026-08

See it in the full glossary →