Reward Model
奖励模型RMCommonA trained scoring network that outputs a reward telling how good a given behavior or outcome is.
A reward model is a trained neural network: it takes a piece of behavior — a robot trajectory video, a large model's answer — as input and outputs a reward score, standing in for a hand-written reward function. A landmark example is Christiano and colleagues' 2017 work, which trained a reward model from human preferences between pairs of trajectories, teaching a simulated robot new behaviors from roughly an hour of human feedback; OpenAI's InstructGPT (2022) put this inside the RLHF (reinforcement learning from human feedback) pipeline, making it a standard step in post-training large models. Many robot tasks resist a precise, hand-written reward (was the shirt actually folded neatly?), and a reward model gives reinforcement learning, failure detection, and data filtering a usable signal instead. Common forms include success detectors, progress reward models, and simply having a VLM score the outcome; the risk is that a policy learns to exploit the model's blind spots — reward hacking.
ExampleRobometer (RSS 2026) trains on RBM-1M, a dataset of over 1 million trajectories including many failures and suboptimal attempts, learning both “task progress at each frame” and “which of two trajectories on the same task is better” at once, producing a general-purpose reward model usable across many robots.
- Also called
- RM, Learned Reward Model
- Related
- Reward Function · Reinforcement Learning from Human Feedback · Progress Reward Model · Success Detector · VLM-as-Reward · Reward Hacking
- Sources
- Deep reinforcement learning from human preferences (Christiano et al., arXiv 1706.03741)
Training language models to follow instructions with human feedback (InstructGPT, arXiv 2203.02155)
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons (arXiv 2603.02115) - As of
- 2026-05