Gradient Accumulation
梯度累积AdvancedAccumulating gradients over several small batches before applying one combined parameter update, simulating a larger batch size.
This is the most common trick for when there isn't enough memory for the batch size you want. Each small batch runs its forward and backward pass as usual, but the optimizer isn't called yet — the gradients simply accumulate on the parameters; only after N small batches have been processed does one parameter update actually happen, after which the gradients are zeroed. The effective batch size becomes the per-GPU batch size times the number of accumulation steps times the number of GPUs, at the cost of fewer updates and more time per update. Two things need care: the loss must be divided by the number of accumulation steps (or normalized by total token count, when computing a token-level loss), or the gradient ends up scaled up too much; and layers like BatchNorm that depend on batch statistics still only ever see the small batch. Training libraries like Hugging Face Accelerate offer this as a ready-made setting, and it's used often when fine-tuning a VLA on a single GPU or just a few.
ExampleIf a single GPU only fits 8 samples but the target batch size is 64, set the accumulation steps to 8: run backward on 8 small batches in a row, then call optimizer.step() once.
- Related
- Batch Size · Gradient Checkpointing (Activation Recomputation) · Mixed-Precision Training · Distributed Training · Optimizer · GPU Memory (VRAM)
- Sources
- Performing gradient accumulation with Accelerate (Hugging Face 文档) (Chinese)