Embodied AI Glossary中文

Bellman Equation

贝尔曼方程Common

A recursive formula that breaks a state's value into the immediate reward plus the discounted value of the next state.

The Bellman equation is named after the American mathematician Richard Bellman, and it comes from the dynamic programming method he introduced. It states that the value of a state equals the immediate reward earned there plus a discount factor times the expected value of the next state. This turns the hard problem of “how good is this in the long run” into a step-by-step recursion. The version written for a fixed policy is called the Bellman expectation equation; the version that takes the maximum over actions is the Bellman optimality equation. Nearly all value-based reinforcement learning builds on this idea — Q-learning, Deep Q-Networks (DQN), and temporal-difference learning all use it to construct their training targets, pushing the network's predictions to satisfy this self-consistent relationship between a state and what follows it.

ExampleQ-learning's training target is r + γ·max Q(s′, a′): when a robot takes action a in state s, gets reward r, and lands in state s′, it updates Q(s, a) toward that target — which is exactly the Bellman optimality equation in use.

Also called
Bellman Optimality Equation, Bellman Expectation Equation
Related
Value Function · Q-Function · Discount Factor · Temporal-Difference Learning · Q-Learning · Markov Decision Process
Sources
Wikipedia: Bellman equation
OpenAI Spinning Up: Key Concepts in RL

See it in the full glossary →