Embodied AI Glossary中文

NVIDIA TensorRT-LLM

TensorRT-LLMAdvanced

NVIDIA's open-source library that speeds up large language model inference on NVIDIA GPUs.

TensorRT-LLM is an open-source large language model inference library that NVIDIA released in 2023, built on top of TensorRT with a Python interface. It packages together the optimizations commonly used for LLM inference: paged management of the KV cache (the stored attention keys and values from previous steps, kept so they don't need recomputing), dynamic batching, quantization formats like FP8 and INT4, speculative decoding, multi-GPU tensor parallelism, and hand-written high-performance kernels for operations such as attention. Users load model weights in a format such as Hugging Face's, and get back an inference service that runs much faster than native PyTorch. In embodied AI, it is often used to accelerate the language-model backbone of a VLA, or a cloud-hosted “brain” service that a robot calls over the network.

ExampleDeploy a Llama model at FP8 precision on an H100 GPU using TensorRT-LLM, as a cloud service for robot task planning.

Also called
TRT-LLM
Related
NVIDIA TensorRT · vLLM · Key-Value Cache · Speculative Decoding · Post-Training Quantization · Inference Deployment
Sources
NVIDIA/TensorRT-LLM (GitHub)
As of
2025

See it in the full glossary →