Bootstrapping (in Reinforcement Learning)
自举AdvancedUpdating a state's estimated value using the agent's own current estimate of the next state's value.
Bootstrapping is a basic technique for updating a value function in reinforcement learning: the update partly relies on an existing value estimate, rather than only on rewards actually received. Temporal-difference (TD) learning is the classic example: when updating the value of a given step, the target is “this step's reward plus the discounted estimated value of the next state,” with no need to wait for the episode to end; by contrast, Monte Carlo methods must run a full episode to completion and use the actual return to update. The benefit of bootstrapping is that it can learn at every step, has low variance, and works for tasks with no natural endpoint; the cost is that the target itself carries bias, and estimation errors can propagate forward. When bootstrapping, function approximation (such as a neural network), and off-policy learning all appear together, training tends to become unstable — this combination is called the “deadly triad,” and DQN's experience replay and target network were both designed to stabilize training against it. This is unrelated to the bootstrap resampling technique in statistics.
ExampleA robot arm opening a drawer gets 0 reward at every step and 1 when the drawer is fully open. When TD learning updates the value of the “gripper is holding the handle” state, it simply takes the current value estimate of the “drawer half-open” state, multiplies it by the discount factor, and uses that as the target — no need to wait for the episode to finish.
- Also called
- Bootstrap
- Related
- Temporal-Difference Learning · Value Function · Bellman Equation · Target Network · Monte Carlo Methods · Overestimation Bias
- Sources
- Lilian Weng: A (Long) Peek into Reinforcement Learning
Wikipedia: Temporal difference learning