Embodied AI Glossary中文

Pruning

剪枝Advanced

Deleting unimportant weights, channels, or layers from a model to make it smaller and faster.

Pruning is a class of model compression methods: it identifies parts of a network that have little effect on the output and removes them. Deleting individual weights is called unstructured pruning; deleting whole channels, attention heads, or layers is called structured pruning. The idea traces back to LeCun and colleagues' 1989 Optimal Brain Damage; in 2015, Song Han and colleagues proposed a three-step 'train, prune small weights, retrain' recipe that cut AlexNet's parameters from 61 million to 6.7 million and compressed VGG-16 13x, with almost no drop in ImageNet accuracy. Unstructured pruning produces a sparse matrix, which needs specialized hardware or libraries to actually speed things up; structured pruning shrinks the matrix directly, making it easier to get a real speedup. In embodied AI, pruning is often combined with quantization and knowledge distillation to fit a large model onto a robot's onboard compute, and VLA-specific work has emerged that prunes redundant language layers or visual tokens.

ExampleEfficientVLA, with no additional training, prunes functionally redundant layers from CogACT's language module, filters out unimportant visual tokens, and caches intermediate features in the diffusion action head, giving a 1.93x inference speedup on SIMPLER with only a 0.6-point drop in success rate.

Also called
Network Pruning
Related
Visual Token Pruning · Post-Training Quantization · Quantization-Aware Training · Knowledge Distillation · On-Device / Edge Deployment · Early Exit
Sources
Learning both Weights and Connections for Efficient Neural Networks (Han et al., arXiv 1506.02626)
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models (arXiv 2506.10100)
As of
2025-06

See it in the full glossary →