Embodied AI Glossary中文

Reward Function

奖励函数Essential

The scoring rule in reinforcement learning that rates an agent's behavior at each step, defining what counts as doing well.

The reward function is a core piece of reinforcement learning (RL): after every action the agent takes, the environment returns a scalar score based on the current state, the action, and the resulting next state, typically written r = R(s, a, s′). The agent's objective is to maximize the sum of these rewards over time (the return), so the reward function effectively defines what the task is. For robot tasks, reward functions are usually hand-written, which is harder than it sounds. Giving credit only on full success (a sparse reward) makes learning slow, but a carelessly designed reward is easy for a policy to exploit — it drives the score up without actually solving the task, a failure mode called reward hacking. This has led to techniques like reward shaping (adding extra intermediate rewards to guide learning), learned reward models, and even having a large language model write reward code automatically.

ExampleA quadruped locomotion reward is usually a weighted sum of terms: a bonus for tracking the target velocity, a penalty for excessive joint torque, and a penalty for falling. NVIDIA and collaborators' Eureka has GPT-4 write reward-function code directly, and it beat reward functions written by human experts on 83% of 29 tasks.

Also called
Reward, Reward Signal
Related
Reinforcement Learning · Return · Reward Shaping · Sparse Reward · Reward Hacking · Eureka
Sources
OpenAI Spinning Up: Key Concepts in RL
Eureka: Human-Level Reward Design via Coding Large Language Models (arXiv 2310.12931)

See it in the full glossary →