Embodied AI Glossary中文

Common Training and Inference GPUs (RTX 4090 / A100 / H100 / B200)

常用 GPU 型号(RTX 4090 / A100 / H100 / B200)Common

The handful of NVIDIA GPUs most often used to train and run embodied-AI models, differing mainly in memory and compute.

These are the NVIDIA GPUs that show up most often in embodied-AI papers and code repositories. The RTX 4090 (Ada architecture, 24 GB of memory) is a consumer gaming card, often used for single-machine inference and small-scale fine-tuning. The A100 (Ampere architecture, 40 or 80 GB) and H100 (Hopper architecture, 80 GB, with FP8 support) are data-center cards, used in multi-GPU clusters for pretraining and full-parameter fine-tuning. The B200 (Blackwell architecture) has up to about 180 GB of memory per card, and an 8-GPU DGX B200 system has roughly 1.4 TB of memory combined. When choosing a card, memory capacity comes first, since it determines how large a model and batch size fit; compute throughput, supported numeric precisions, and inter-GPU interconnect bandwidth matter next. Models running on the robot itself typically run on an edge chip like a Jetson rather than on any of these cards.

ExampleThe openpi repository's guidance: running π0 inference needs at least 8 GB of memory and LoRA fine-tuning needs at least 22.5 GB, both of which an RTX 4090 can handle; full-parameter fine-tuning needs at least 70 GB, requiring an 80 GB A100 or an H100.

Also called
RTX 4090, A100, H100, B200
Related
GPU Memory (VRAM) · NVIDIA Jetson · LoRA · Full Fine-Tuning · Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4) · AutoDL
Sources
openpi README(Hardware Requirements)
NVIDIA DGX B200 User Guide: Introduction
As of
2026-09

See it in the full glossary →