Multi-Task Learning
多任务学习MTLCommonTraining one model on several related tasks at once, so it shares knowledge and each task helps the others.
Rich Caruana's 1997 paper “Multitask Learning” laid out this idea systematically: train several related tasks in parallel sharing one representation, so what one task learns helps the others learn better, with each task's training signal acting as an extra inductive bias (a model's built-in leaning toward what kind of solution is reasonable) for the rest. The most common deep-learning implementation is hard parameter sharing: several tasks share one backbone network, each with its own output head. In robotics, a multi-task policy is told what to do right now via a language instruction or a goal image, and one network covers many skills — opening a drawer, grasping, placing — and essentially every VLA today is a multi-task model. The difficulty is that tasks can interfere with each other (negative transfer), which requires tuning the data mix and network architecture.
ExampleGoogle's RT-1 used 13 robots over 17 months to collect more than 130,000 demonstrations spanning over 700 language instructions, with one Transformer policy learning all of them and reaching 97% success on instructions it had seen before.
- Also called
- MTL, Multitask Learning
- Related
- Positive / Negative Transfer · Transfer Learning · Language-conditioned Policy · Generalist Policy · Data Mixture · RT-1
- Sources
- Caruana (1997), Multitask Learning, Machine Learning 28
An Overview of Multi-Task Learning in Deep Neural Networks (Ruder, 2017)
RT-1: Robotics Transformer for Real-World Control at Scale (project page)