Embodied AI Glossary中文

Backbone Freezing

冻结骨干网络Common

Keeping a pretrained backbone network's parameters fixed during training and updating only the newly added parts.

A backbone is the part of a model responsible for extracting features, usually a pretrained network such as a ResNet, a ViT, or an entire vision-language model (VLM). Freezing it means those parameters are not updated during training (in PyTorch, by setting requires_grad to False), and only newly added output heads or action modules are trained. The upside is lower memory use, faster training, less overfitting when data is scarce, and less catastrophic forgetting; the downside is that the backbone can't learn features the new task might need. The trade-off shows up concretely in VLA models: OpenVLA found that freezing the vision encoder loses fine-grained spatial information and hurts control performance, while NVIDIA's GR00T N1 freezes the VLM's language component during both pretraining and post-training and trains only the remaining modules.

ExamplePyTorch's own transfer-learning tutorial fine-tunes an ImageNet-pretrained ResNet-18 to classify ants versus bees: every layer except the last is frozen, and only the newly swapped-in fully connected layer is trained.

Also called
Freezing the Backbone, Frozen Backbone, Frozen Pretrained Layers
Related
Backbone Network · Fine-tuning · Parameter-Efficient Fine-Tuning · Catastrophic Forgetting · Knowledge Insulation · Stop-Gradient
Sources
PyTorch Tutorial: Transfer Learning for Computer Vision
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)
As of
2025-03

See it in the full glossary →