Checkpoint
检查点ckptEssentialA saved snapshot of a model's parameters, taken during or after training, that can be loaded to run inference or resume training.
A checkpoint is a saved snapshot of a model's parameters at a given moment. Training a large model often takes days or even weeks, so a program will save the parameters to a file every so often, both to avoid starting over after a crash and to make it easy to later pick the version that performed best on validation. If a checkpoint is only meant for inference, saving the model weights is enough; to resume training exactly where it left off, the optimizer state, current epoch, and similar information need to be saved alongside it — PyTorch's documentation notes that this kind of full checkpoint is typically 2 to 3 times the size of the weights alone. Common file formats include .pt, .pth, .ckpt, and safetensors. When an open-source VLA project says it has “released weights,” it means it has published a checkpoint, which anyone can download to run inference on directly or continue fine-tuning from.
ExamplePhysical Intelligence's openpi repository provides base checkpoints such as pi0_base and pi05_base for fine-tuning, as well as already fine-tuned checkpoints like pi05_libero and pi0_aloha_towel that can be downloaded and run for inference right away.
- Also called
- Model Weights, Model Checkpoint, ckpt
- Related
- Fine-tuning · Pre-training · Open-weight Model · safetensors · Inference Deployment · Epoch
- Sources
- PyTorch Tutorials: Saving and Loading Models
Google Machine Learning Glossary: checkpoint
GitHub: Physical-Intelligence/openpi