Embodied AI Glossary中文

Twin Delayed DDPG

双延迟深度确定性策略梯度TD3Common

A continuous-action reinforcement-learning algorithm that adds three fixes to DDPG specifically to curb Q-value overestimation.

TD3 was introduced by Scott Fujimoto, Herke van Hoof, and David Meger at ICML 2018, as an improved version of DDPG (an actor-critic algorithm that outputs deterministic actions). It's off-policy (able to reuse old data repeatedly) and applies only to continuous actions. DDPG's Q-network tends to overestimate action values, and the policy learns to exploit those inflated estimates, which often makes training unstable. TD3 addresses this with three changes: training two Q-networks and taking the smaller one when computing the target (twin); updating the policy and target networks only once for every two Q-network updates (delayed); and adding clipped noise to the target action, smoothing how the Q-value changes with the action. TD3 and SAC are the two most common baselines for continuous control, and the offline reinforcement-learning method TD3+BC is also built on it.

ExampleThe original paper tests on OpenAI Gym's MuJoCo continuous-control tasks (such as HalfCheetah, Hopper, and Walker2d), where TD3 outperformed DDPG and other leading algorithms of the time.

Also called
TD3, Twin Delayed Deep Deterministic Policy Gradient
Related
Deep Deterministic Policy Gradient · Overestimation Bias · Target Network · Soft Actor-Critic · Off-Policy · Value Function
Sources
Addressing Function Approximation Error in Actor-Critic Methods (TD3, arXiv 1802.09477)
OpenAI Spinning Up: Twin Delayed DDPG
A Minimalist Approach to Offline Reinforcement Learning (TD3+BC, arXiv 2106.06860)

See it in the full glossary →