Embodied AI Glossary中文

Sim-to-Real Correlation

仿真-真机相关性Common

A measure of whether simulated evaluation scores actually reflect real-robot performance, usually reported as a Pearson correlation coefficient r.

Before using simulation for evaluation, it needs to be confirmed that a policy which scores well in simulation is also strong on the real robot. A common approach picks a set of policies, measures their success rate in both simulation and on the real robot, and computes the Pearson correlation coefficient r between the two sets of numbers (ranging from -1 to 1, with values closer to 1 meaning better agreement). The 2024 SIMPLER (SimplerEnv) paper ran this comparison for policies such as RT-1, RT-1-X, and Octo on the Google Robot and WidowX platforms, and pointed out that r only captures linear fit and can miss cases where the ranking itself gets flipped, so it also proposed Maximum Mean Rank Violation (MMRV, where lower is better). When correlation is high, a simulation benchmark can substitute for expensive real-robot testing when picking a model; when it is low, gains seen in simulation may just reflect overfitting to the simulator.

ExampleSIMPLER narrows the gap through visual matching (compositing a real background onto a green screen, aligning textures) and system identification of control parameters, demonstrating a strong correlation between simulated and real-robot scores across about 1,500 evaluation episodes.

Also called
Pearson r, Sim-Real Alignment
Related
Simulation-Based Evaluation · Real-World Evaluation · Mean Maximum Rank Violation · SimplerEnv · Visual Matching (SimplerEnv) · Sim-to-Real Gap (Reality Gap)
Sources
Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)
SIMPLER 项目页 (Chinese)
As of
2024-05

See it in the full glossary →