Training from Scratch
从零训练CommonTraining a model directly on the target data with randomly initialized parameters, borrowing no pretrained weights at all.
Training from scratch means starting a network's weights from random values and training only on data for the current task, as opposed to loading a pretrained model and fine-tuning it. The benefit is not being constrained by the pretraining data's distribution or the pretrained model's architecture, and a simpler pipeline; the downside is needing more data and training time, and being more prone to overfitting when data is scarce. Kaiming He and colleagues' 2018 paper “Rethinking ImageNet Pre-training” found that for object detection, training from scratch for long enough can match ImageNet pretraining, showing that one of pretraining's main benefits is simply faster convergence. Because robot data is scarce, papers often use training from scratch as a baseline, to measure how much benefit a pretrained representation or large-scale pretraining actually provides — though pretraining isn't always better: Diffusion Policy's visual encoder, for instance, is an un-pretrained ResNet-18 trained end-to-end from scratch.
ExampleThe R3M paper compares across 12 simulated manipulation tasks: a visual representation pretrained on Ego4D human video beats a visual encoder trained from scratch by more than 20 percentage points in success rate.
- Also called
- Random Initialization Training, From Scratch
- Related
- Pre-training · Fine-tuning · Baseline · Overfitting · Transfer Learning · Pre-trained Visual Representation
- Sources
- Rethinking ImageNet Pre-training (He et al., arXiv 1811.08883)
R3M: A Universal Visual Representation for Robot Manipulation (arXiv 2203.12601)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)