Hugging Face Accelerate
Accelerate 库AdvancedHugging Face's training-helper library that lets the same PyTorch code run across multiple GPUs with minimal changes.
Accelerate is an open-source PyTorch helper library from Hugging Face. A training script written for a single GPU only needs a few added lines — creating an Accelerator and wrapping the model, optimizer, and data loader with it — to run unchanged across a single GPU, multiple GPUs, multiple machines, or with mixed precision, and it can switch to a distributed backend such as DeepSpeed or FSDP (Fully Sharded Data Parallel) without the developer having to write process initialization or gradient synchronization by hand. The companion commands accelerate config and accelerate launch generate a configuration and then launch training. Fine-tuning a large model like a VLA often needs multiple GPUs, and a lot of open-source robot-learning code uses Accelerate to manage that distributed training.
ExampleRunning accelerate launch starts a VLA fine-tuning script across 8 GPUs, with bf16 mixed precision turned on.
- Also called
- accelerate
- Related
- Distributed Data Parallel (DDP) · Fully Sharded Data Parallel (FSDP) · DeepSpeed · Mixed-Precision Training · PyTorch · Hugging Face
- Sources
- Accelerate 官方文档 (Chinese)