Pre-training
预训练EssentialTraining a foundation model on massive general-purpose data first, before adapting it to any specific task.
Pretraining is the first stage of building a foundation model: the model is trained on data that's both large in scale and broad in coverage — web text, image-text pairs, video, manipulation data from many different robots — so it learns general representations and capabilities. The result is called a pretrained model, or base model, which is later adapted to a specific downstream task through fine-tuning or post-training. Its value is in doing the expensive part of general learning once, so that downstream tasks only need a small amount of additional data. In embodied AI, pretraining typically happens in two layers: a VLA model first inherits a vision-language model that was already pretrained on internet image-text data, then continues training on large-scale data from many different robots — π0, for example, used over 10,000 hours of robot data for this second layer.
ExampleOpenVLA (7B parameters), built on a Llama 2 language model plus DINOv2 and SigLIP vision encoders, was pretrained on 970,000 real-robot demonstrations from Open X-Embodiment, and can then be fine-tuned to a new task with LoRA on a consumer GPU.
- Also called
- Pretraining
- Related
- Post-training · Fine-tuning · Foundation Model · Downstream Task · Scaling Law · Self-Supervised Learning
- Sources
- Google Machine Learning Glossary: pre-trained model
Kim et al. 2024: OpenVLA: An Open-Source Vision-Language-Action Model
Black et al. 2024: π0: A Vision-Language-Action Flow Model for General Robot Control