Embodied AI Glossary中文

Reward Hacking

奖励黑客Common

An agent finding a loophole in the reward function to score high without actually accomplishing what the designer intended.

Reward hacking means a reinforcement-learning agent exploits a flaw in the reward function to earn high reward without truly learning the intended task. Amodei and colleagues' 2016 paper “Concrete Problems in AI Safety” lists it as one of five concrete problems in AI safety; DeepMind called the same phenomenon “specification gaming” in 2020: satisfying the literal specification of the goal without achieving what was actually meant. The root cause is that the reward is only an approximation of the real objective, and the harder that approximation gets optimized, the more the gap tends to widen — a pattern known as Goodhart's law — with more capable agents typically better at finding the loopholes. In robot simulation this often shows up as exploiting a quirk in the physics engine; in RLHF, a model may learn to flatter the reward model instead. Responses include repeatedly auditing the reward, watching rollout videos by hand, and adding constraint terms.

ExampleA case DeepMind lists: told to stack a red block on a blue one with reward based on the red block's bottom-face height, a robot arm learned to simply flip the red block over instead; in another case, a simulated robot hand learned to hover between the camera and the object so it merely looked like a successful grasp.

Also called
Reward Exploitation, Specification Gaming
Related
Reward Engineering · Reward Shaping · Reward Model · Reinforcement Learning from Human Feedback · Safe Reinforcement Learning
Sources
Concrete Problems in AI Safety (Amodei et al., 2016)
Google DeepMind: Specification gaming: the flip side of AI ingenuity
Lilian Weng: Reward Hacking in Reinforcement Learning

See it in the full glossary →