GPU Memory (VRAM)
GPU 显存VRAMCommonThe high-speed memory built into a graphics card that limits how big a model and batch can be.
GPU memory (VRAM) is the dedicated high-speed memory built into a graphics card. Model parameters, gradients, optimizer state, intermediate activations, and input data all have to fit in it for the GPU to compute anything. How much VRAM you have directly sets the largest model you can train or run, and the largest batch size (how many samples you feed in at once) you can use. Training needs far more VRAM than inference, because it also has to store gradients and optimizer state. Common fixes when you run out include lower numerical precision (e.g., BF16, INT8 quantization), LoRA (which trains only a small number of parameters), gradient checkpointing, and splitting the model across multiple GPUs. For embodied AI, VRAM determines whether a VLA (vision-language-action) model can be fine-tuned on a lab's GPUs, and whether it can be squeezed onto an onboard chip on the robot itself.
ExampleThe openpi documentation gives these reference numbers: inference for π0 needs over 8 GB of VRAM, LoRA fine-tuning needs roughly 22.5 GB or more, and full fine-tuning needs roughly 70 GB or more.
- Also called
- VRAM, Video RAM
- Related
- Parameter Count (Model Size) · LoRA · Post-Training Quantization · Gradient Checkpointing (Activation Recomputation) · Batch Size · On-Device / Edge Deployment
- Sources
- Video random-access memory - Wikipedia
openpi - Physical Intelligence (GitHub)