Post-Training Quantization
训练后量化PTQAdvancedConverting a trained model's weights and activations directly to low-bit numbers after training, with no retraining needed.
Post-training quantization is a model compression method: after a model is trained at normal precision (FP32, BF16, etc.), its weights, and sometimes its activations too, are converted to a low-bit representation such as INT8 or INT4. It usually only needs a small batch of calibration data to measure the value ranges, requires no labeled data, and involves no retraining. Qualcomm AI Research's quantization white paper summarizes that most models keep accuracy close to floating point at 8 bits with PTQ; pushing lower tends to hurt accuracy and needs methods like GPTQ, which use second-order information, or a switch to quantization-aware training. For embodied AI, it mainly serves deployment: VLAs routinely have billions of parameters, while a robot's onboard memory and compute are limited, and quantization directly cuts memory use and speeds up inference. Deployment tools such as TensorRT and ONNX Runtime both support PTQ.
ExampleThe OpenVLA paper runs its 7B model at 4-bit quantized inference, reaching 71.9% success on BridgeData V2 tasks versus 71.3% at BF16, while GPU memory drops from 16.8GB to 7.0GB.
- Also called
- PTQ, Offline Quantization
- Related
- Quantization-Aware Training · Pruning · On-Device / Edge Deployment · Inference Latency · NVIDIA TensorRT · Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4)
- Sources
- A White Paper on Neural Network Quantization (Nagel et al., arXiv 2106.08295)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (arXiv 2210.17323) - As of
- 2024-06