Twin Delayed DDPG
双延迟深度确定性策略梯度TD3CommonA continuous-action reinforcement-learning algorithm that adds three fixes to DDPG specifically to curb Q-value overestimation.
TD3 was introduced by Scott Fujimoto, Herke van Hoof, and David Meger at ICML 2018, as an improved version of DDPG (an actor-critic algorithm that outputs deterministic actions). It's off-policy (able to reuse old data repeatedly) and applies only to continuous actions. DDPG's Q-network tends to overestimate action values, and the policy learns to exploit those inflated estimates, which often makes training unstable. TD3 addresses this with three changes: training two Q-networks and taking the smaller one when computing the target (twin); updating the policy and target networks only once for every two Q-network updates (delayed); and adding clipped noise to the target action, smoothing how the Q-value changes with the action. TD3 and SAC are the two most common baselines for continuous control, and the offline reinforcement-learning method TD3+BC is also built on it.
ExampleThe original paper tests on OpenAI Gym's MuJoCo continuous-control tasks (such as HalfCheetah, Hopper, and Walker2d), where TD3 outperformed DDPG and other leading algorithms of the time.
- Also called
- TD3, Twin Delayed Deep Deterministic Policy Gradient
- Related
- Deep Deterministic Policy Gradient · Overestimation Bias · Target Network · Soft Actor-Critic · Off-Policy · Value Function
- Sources
- Addressing Function Approximation Error in Actor-Critic Methods (TD3, arXiv 1802.09477)
OpenAI Spinning Up: Twin Delayed DDPG
A Minimalist Approach to Offline Reinforcement Learning (TD3+BC, arXiv 2106.06860)