Embodied AI Glossary中文

Reward Shaping

奖励塑形Common

Adding intermediate guidance rewards on top of a task's original reward, so the agent learns the task faster.

Reward shaping means adding extra reward for intermediate progress on top of the task's own reward (for instance, only +1 on success), such as scoring higher the closer a robot arm gets to its target. It mainly addresses the problem that under a sparse reward, an agent can go a very long time with no feedback and simply fail to learn. The risk is that shaping it wrong changes what the “optimal behavior” actually is: a classic example is a simulated cyclist that learned to circle endlessly near a goal because getting closer was rewarded and moving away wasn't penalized. Ng, Harada, and Russell proved in a 1999 ICML paper that as long as the extra reward is written as the difference of a potential function, F = γΦ(s′) − Φ(s) (Φ scores each state, γ is the discount factor), the optimal policy is provably unchanged — this is called potential-based reward shaping. In embodied AI, a legged-locomotion reward is usually a weighted combination of velocity tracking, posture, energy use, and foot airtime, which is essentially a large amount of hand-tuned shaping; projects like Eureka try to have a large model write this kind of reward code automatically.

ExampleTraining a robot arm to push a block to a target: the raw reward gives +1 only when the block arrives; after shaping, each step also rewards “how much closer did the gripper get to the block” and “how much closer did the block get to the target,” giving the policy a learning signal from early in training.

Also called
Potential-based Reward Shaping
Related
Sparse Reward · Dense Reward · Reward Function · Reward Engineering · Reward Hacking · Eureka
Sources
Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping (Ng, Harada, Russell, ICML 1999)
Reward Hacking in Reinforcement Learning (Lilian Weng, 2024)

See it in the full glossary →