Q-Transformer
AdvancedA large-scale offline reinforcement learning method that uses a Transformer to estimate Q-values one action dimension at a time.
Q-Transformer was released by Yevgen Chebotar, Sergey Levine, and colleagues at Google DeepMind in September 2023 and published at CoRL 2023. The problem it targets: robots often have both human demonstrations and a much larger pool of autonomously collected data, including failures, and offline reinforcement learning (training only from existing data, without further interaction) offers a way to use both. The method represents the Q-function (a value function that scores a state-action pair) with a Transformer: each action dimension is discretized into a number of bins, and the Q-value for each dimension is predicted one at a time, similar to generating tokens. Training is stabilized with conservative regularization (pushing the value of unseen actions down to a minimum) and Monte Carlo return estimates. On RT-1's real-robot multi-task benchmark, Q-Transformer outperforms RT-1, IQL, and Decision Transformer, and is especially good at making use of failure data.
ExampleOn RT-1's real-robot multi-task data, Q-Transformer trains on both human demonstrations and failure episodes from the robot's own autonomous attempts, outperforming RT-1, which only imitates successful demonstrations.
- Also called
- Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions
- Related
- QT-Opt · Offline Reinforcement Learning · Q-Function · RT-1 · Conservative Q-Learning · Decision Transformer
- Sources
- Q-Transformer (arXiv 2309.10150)
Q-Transformer project page - As of
- 2023-09