VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
VLA-ArenaAdvancedPeking University's open-source VLA evaluation framework, grading capability boundaries across task, language, and vision axes.
VLA-Arena is an open-source VLA (vision-language-action model) evaluation framework released by Peking University's PKU-Alignment team, posted to arXiv in December 2025 and later accepted at ICML 2026, with simulation built on LIBERO and robosuite. It breaks difficulty into three independent axes: task structure, language instructions, and visual observations. There are 11 task suites totaling 170 tasks, split into four categories — safety, distractors, extrapolation, and long-horizon — with each suite graded L0–L2, where fine-tuning is only allowed on L0; language (W0–W4) and vision (V0–V4) perturbations can then be layered onto any task. Using this setup, the authors find that current VLAs tend to memorize their training tasks, understand vision only shallowly, and often disregard safety constraints.
ExampleThe paper found that a model ranking near the top on L0 could be overtaken by other models once moved to L1 or L2; the authors treat this kind of rank reversal as evidence that the three difficulty levels each carry independent information.
- Also called
- VLA-Arena Benchmark
- Related
- Vision-Language-Action Model · LIBERO Benchmark · robosuite · Generalization / Robustness Evaluation · Embodied Safety · Benchmark
- Sources
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models (arXiv 2512.22539)
VLA-Arena 项目主页 (Chinese) - As of
- 2026-08