Temporal-Difference Learning
时序差分学习TDCommonUpdating a value estimate using “this step's reward plus the estimated value of the next state,” without waiting for the episode to end.
Temporal-difference (TD) learning was systematically introduced by Richard Sutton in a 1988 paper and is the core method for estimating a value function (how much total return follows from a given state or action) in reinforcement learning. It doesn't need to wait for an episode to finish to get the true return; instead it updates after every single step, treating “the reward r actually received, plus the discounted estimated value of the next state, γV(s′)” as a target, calling the gap between that target and the current estimate V(s) the TD error, and nudging V(s) toward the target by that amount. This “using an estimate to update an estimate” approach is called bootstrapping, and it has lower variance than Monte Carlo methods, which wait for the full episode, letting it learn while still acting — at the cost of introducing some bias. Q-learning, DQN, and the critic networks in SAC and TD3 are all trained with TD targets; the early landmark system TD-Gammon used it to reach expert-level backgammon play.
ExampleA robot arm takes one step and gets a reward of 0; the critic estimates the next state's value at 0.8, and with discount factor γ = 0.99, the TD target is 0 + 0.99 × 0.8 = 0.792. If the current state's estimated value is 0.5, the TD error is 0.292, and the network nudges its estimate toward 0.792.
- Also called
- TD Learning, TD Error
- Related
- Value Function · Bellman Equation · Bootstrapping (in Reinforcement Learning) · Q-Learning · Monte Carlo Methods · Generalized Advantage Estimation
- Sources
- Temporal difference learning (Wikipedia)
Soft Actor-Critic (OpenAI Spinning Up)