Bellman Equation
贝尔曼方程CommonA recursive formula that breaks a state's value into the immediate reward plus the discounted value of the next state.
The Bellman equation is named after the American mathematician Richard Bellman, and it comes from the dynamic programming method he introduced. It states that the value of a state equals the immediate reward earned there plus a discount factor times the expected value of the next state. This turns the hard problem of “how good is this in the long run” into a step-by-step recursion. The version written for a fixed policy is called the Bellman expectation equation; the version that takes the maximum over actions is the Bellman optimality equation. Nearly all value-based reinforcement learning builds on this idea — Q-learning, Deep Q-Networks (DQN), and temporal-difference learning all use it to construct their training targets, pushing the network's predictions to satisfy this self-consistent relationship between a state and what follows it.
ExampleQ-learning's training target is r + γ·max Q(s′, a′): when a robot takes action a in state s, gets reward r, and lands in state s′, it updates Q(s, a) toward that target — which is exactly the Bellman optimality equation in use.
- Also called
- Bellman Optimality Equation, Bellman Expectation Equation
- Related
- Value Function · Q-Function · Discount Factor · Temporal-Difference Learning · Q-Learning · Markov Decision Process
- Sources
- Wikipedia: Bellman equation
OpenAI Spinning Up: Key Concepts in RL