Success Rate
成功率SREssentialThe fraction of attempts at the same task that a policy completes successfully — the most common robotics metric.
Success rate is the most commonly used evaluation metric in robot learning: a policy attempts the same task N times (each attempt called an episode, usually with a randomized initial object position), the number of successes is counted against a predefined criterion (such as “the block ends up in the target zone”), and that count is divided by N. It's intuitive and easy to compare across methods, but it carries limited information — it only looks at the final outcome, and doesn't distinguish a near-miss from a total failure — so it's often reported alongside subtask success rate, a progress score, or average completed length. Success rate is also very sensitive to the number of trials: one extra success out of 25 trials shifts the number by 4 percentage points, so rigorous papers report the number of trials and random seeds used, and ideally a confidence interval. Simulation can run hundreds or thousands of episodes at once; real-robot testing is usually limited to a few dozen.
ExampleThe ACT paper averages success rate over 3 random seeds with 50 trials each for simulated tasks; for real-robot tasks it runs 25 trials each, with “open a sealed bag” reaching 88% success.
- Also called
- SR, Task Success Rate
- Related
- Progress Score · Average Length (CALVIN) · Episode · Simulation-Based Evaluation · Real-World Evaluation · Statistical Rigor in Policy Evaluation (Confidence Intervals / Sequential Testing / Multiple Seeds)
- Sources
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)
Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER)