Embodied AI Glossary中文

Decision Transformer

决策 TransformerDTAdvanced

Treats reinforcement learning as sequence modeling: given a target return, a GPT-style model predicts actions one step at a time.

Decision Transformer was proposed in June 2021 by Lili Chen, Kevin Lu, and colleagues at Berkeley, Facebook AI Research, and Google Brain, published at NeurIPS 2021. It doesn't fit a value function or compute policy gradients; instead, it writes a trajectory as a sequence of tokens alternating “return-to-go” (how much reward is still wanted from this point to the end), state, and action, and trains a GPT-style Transformer with causal masking, using supervised learning, to predict the next action. At test time, a target return is set at the start, and the reward received at each step is subtracted from it as the episode goes on. Using only offline data, it matched or beat the leading model-free offline reinforcement-learning methods of the time on Atari, OpenAI Gym, and Key-to-Door tasks, helping make “treat decision-making as sequence modeling” one of the mainstream approaches.

ExampleIn an Atari game, a fairly high target return is set at the start; the model reads in the most recent frames, past actions, and remaining return-to-go, and outputs the next button press; every time points are scored, the remaining return-to-go drops accordingly before the next action is generated.

Also called
DT
Related
Offline Reinforcement Learning · Transformer · Return Conditioning · Causal Attention · Decision Diffuser · Gato
Sources
Decision Transformer (arXiv:2106.01345)
Decision Transformer 项目主页 (Chinese)
As of
2021-06

See it in the full glossary →