Embodied AI Glossary中文

TD-MPC2

Advanced

A reinforcement learning algorithm that plans inside a learned latent-space world model, using one set of hyperparameters across hundreds of control tasks.

TD-MPC2 was released by Nicklas Hansen, Hao Su, and Xiaolong Wang at UC San Diego in October 2023, an ICLR 2024 Spotlight paper and an improvement on the same team's earlier TD-MPC. It is a model-based reinforcement learning method: it first learns a world model that predicts the next state, reward, and value purely in a latent space, with no image reconstruction (that is, no decoder); at decision time, it performs model predictive control (MPC) within that latent space, relying on short-horizon rollouts from the model while longer-horizon returns are estimated by a value function learned via temporal difference (TD) learning, which is where the name comes from. TD-MPC2 performs reliably with the same set of hyperparameters across 104 continuous-control tasks spanning four domains — DMControl, Meta-World, ManiSkill2, and MyoSuite — and the team also trained a 317-million-parameter multi-task agent that handles 80 tasks, showing that capability keeps growing with model and data scale.

ExampleThe same TD-MPC2 multi-task model can control the cheetah-run task in DMControl and also open a drawer with a robot arm in Meta-World, without needing separate hyperparameter tuning for each task.

Also called
TD-MPC 2, TD-MPC2: Scalable, Robust World Models for Continuous Control
Related
Model-Based Reinforcement Learning · World Model · Model Predictive Control · Latent World Model · Temporal-Difference Learning · DreamerV3
Sources
TD-MPC2: Scalable, Robust World Models for Continuous Control (arXiv 2310.16828)
TD-MPC2 project page
As of
2024-01

See it in the full glossary →