Embodied AI Glossary中文

Variant Aggregation (SimplerEnv)

变体聚合VAAdvanced

One of SimplerEnv's two evaluation modes: testing a policy across many visually randomized scene variants and averaging the results.

Variant Aggregation is one of two “real-to-sim” evaluation approaches proposed by SimplerEnv (SIMPLER), from a paper by Xuanlin Li and colleagues at CoRL 2024. Simulated images and real-robot images always differ somewhat; the other approach, Visual Matching, tries to make the simulated image look as close to reality as possible. Variant Aggregation instead does the opposite: it applies heavy visual randomization to a scene, generating multiple environment variants along axes such as background, lighting, distractor objects, table texture, and camera pose, measures success rate separately in each, then averages them to estimate how a policy performs overall under visual variation. The paper uses Mean Maximum Rank Violation (MMRV) and the Pearson correlation coefficient to measure whether simulated results agree with real-robot results; in the paper's experiments, Visual Matching agreed with real robots better overall. VLA papers reporting SimplerEnv results on Google-robot tasks commonly list both a VM and a VA column.

ExampleFor the “pick up the Coke can” task, run the same policy through several variants — a different background, different lighting, added distractors, a different table texture, a different camera angle — and average the success rates across them to get its VA score.

Also called
VA
Related
SimplerEnv · Visual Matching (SimplerEnv) · Mean Maximum Rank Violation · Sim-to-Real Correlation · Visual Randomization · Simulation-Based Evaluation
Sources
Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)
simpler-env/SimplerEnv GitHub 仓库 (Chinese)
As of
2024-05

See it in the full glossary →