Embodied AI Glossary中文

Robometer

Advanced

A general-purpose robot reward model trained on both task-progress labels and pairwise trajectory comparisons.

Robometer was released in March 2026 by researchers from the University of Southern California, the University of Washington, MIT, the Allen Institute for AI, NVIDIA, and other institutions, accepted at RSS 2026. Both reinforcement learning and data curation need a reward model that can judge how well a robot is doing, but earlier approaches relied mainly on frame-by-frame progress labels on expert demonstrations, leaving large amounts of failed and suboptimal trajectories unused. Robometer combines two supervision signals: a frame-level progress loss that anchors the reward's scale using expert data, and a preference loss from pairwise trajectory comparisons that learns which of two trajectories is better, which lets it also learn from failure data. The authors built RBM-1M, a dataset of more than one million trajectories spanning 21 robot embodiments, and the model is based on Qwen3-VL (main version 4B). The resulting reward can be used for online and offline reinforcement learning, failure detection, and retrieving data for imitation learning.

ExampleGiven a video of a robot performing 'put the cup in the drawer,' Robometer outputs a frame-by-frame task-progress score; this curve can be used directly as a dense reward for reinforcement learning, or to judge whether the attempt failed.

Also called
RBM-1M, Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
Related
Reward Model · Progress Reward Model · VLM-as-Reward · Dense Reward · Success Detector · Failure Data
Sources
arXiv 2603.02115: Robometer
Robometer 项目主页 (Chinese)
As of
2026-05

See it in the full glossary →