Foundation Model
基础模型EssentialA large model pretrained on massive data that can later adapt to many different downstream tasks.
The term “foundation model” was coined by more than a hundred Stanford researchers in a long 2021 survey: a model trained on large-scale, diverse data, usually with self-supervised learning, which builds training targets from the data itself rather than from human labels, and that can then adapt to many downstream tasks through fine-tuning or prompting; the paper's examples are BERT, GPT-3, and DALL-E. It changed the old habit of training one model per task: spend a lot of compute once on a general-purpose base, then adjust it a little for each task, at the cost that any flaw in the base gets inherited by every downstream model. In embodied AI, VLAs are usually built on a vision-language model base, and a model pretrained on large amounts of robot data that can adapt to many robots and tasks is commonly called a robot foundation model or embodied foundation model.
Exampleπ0 is built on PaliGemma, Google's open-source vision-language model with 3 billion parameters, adding a 300-million-parameter action expert, and is trained on robot data into a robot foundation model that can then be fine-tuned further for specific tasks.
- Also called
- Foundation Models
- Related
- Large Language Model · Vision-Language Model · Embodied Foundation Model · Pre-training · Fine-tuning · Vision-Language-Action Model
- Sources
- On the Opportunities and Risks of Foundation Models (Bommasani et al., arXiv:2108.07258)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164) - As of
- 2024-10