Embodied AI Glossary中文

Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4)

数值精度格式(FP32 / FP16 / BF16 / FP8 / INT8 / INT4)Common

The bit width and format used to store a model's numbers, which sets its memory use, speed, and accuracy.

Every parameter and intermediate result in a neural network is a number, and how many bits are used to store it directly affects memory, speed, and precision. FP32 is 32-bit single precision, the most stable but the most memory-hungry; FP16 half precision uses only 16 bits with a small numeric range, so training with it easily overflows and needs loss scaling; BF16 also uses 16 bits but keeps the same exponent range as FP32 — wide range, lower precision — and is the default choice for large-model training today; FP8 uses only 8 bits and needs GPUs from the H100 generation or later; INT8 and INT4 are integer formats mainly used to compress models for inference. As a rough rule, memory equals parameter count times bytes per number (4 for FP32, 2 for BF16, 1 for INT8, 0.5 for INT4), so a 7B-parameter model stored in BF16 needs roughly 14 GB, or about 3.5 GB quantized to INT4 — which determines whether a VLA can fit on a robot's onboard compute.

ExampleA 7B model like OpenVLA takes roughly 15 GB of GPU memory loaded in BF16; quantizing it to 4 bits lets it run on smaller GPUs, though action accuracy may drop.

Also called
FP32, FP16, BF16, FP8, INT8, INT4, half precision, single precision
Related
Mixed-Precision Training · Post-Training Quantization · Quantization-Aware Training · GPU Memory (VRAM) · On-Device / Edge Deployment · NVIDIA TensorRT
Sources
Wikipedia: bfloat16 floating-point format
NVIDIA Transformer Engine: Using FP8

See it in the full glossary →