MT-Opt
AdvancedA Google multi-task reinforcement-learning system that used 7 real robots to gather 9,600 hours of data while learning 12 tasks at once.
MT-Opt was released in April 2021 by Dmitry Kalashnikov, Chelsea Finn, Sergey Levine, Karol Hausman, and colleagues at Robotics at Google, a multi-task extension of QT-Opt (Google's large-scale Q-learning-based grasping reinforcement-learning system). The team used 7 robots to continuously gather about 9,600 robot-hours across more than 800,000 episodes over 57 days, learning 12 real tasks at once — picking a specific object, placing it in a container, aligning objects, covering, and more. Two design choices are key: a multi-task success detector that automatically judges whether a task is done and assigns reward, and sharing episodes from one task with the others while rebalancing the amount of data, so tasks with little data can borrow experience from more common ones. Average success rate on rare tasks rose from 1% under single-task QT-Opt to 50%, and a new task could be fine-tuned in about a day. It is an important piece of Google's exploration of scaling up real-robot learning before RT-1.
ExampleTo teach the robot a new task, “cover an object with a towel,” there's no need to train from scratch; collecting about a day's worth of additional data to fine-tune on top of already-learned skills like grasping is enough.
- Also called
- Continuous Multi-Task Robotic Reinforcement Learning at Scale
- Related
- QT-Opt · Reinforcement Learning · Multi-Task Learning · Success Detector · Real-World Reinforcement Learning · Offline Reinforcement Learning
- Sources
- MT-Opt (arXiv 2104.08212)
Multi-Task Robotic Reinforcement Learning at Scale (Google Research blog, 2021-04-19)
MT-Opt project page - As of
- 2021-04