Embodied AI Glossary中文

NVIDIA TensorRT

TensorRTTRTCommon

NVIDIA's inference-acceleration library that compiles a trained model into a faster GPU runtime engine.

TensorRT is NVIDIA's deep-learning inference optimizer and runtime. It reads in a model in a format such as ONNX, performs layer fusion, picks the fastest available GPU kernels, and lowers numerical precision (FP16, INT8, FP8, and so on) to produce an inference “engine” file compiled for a specific GPU. A robot policy has to produce an action within tens of milliseconds, and running it directly in PyTorch is often too slow, so compiling with TensorRT typically cuts latency noticeably — making it a standard step for both Jetson and server-side deployment. Note that an engine is tied to a specific GPU model and TensorRT version, so switching hardware requires recompiling; large language models have their own dedicated variant, TensorRT-LLM.

ExampleExport a trained diffusion policy to ONNX, then compile it into an FP16 engine with trtexec to cut per-step inference latency on a Jetson Orin.

Also called
TRT
Related
Open Neural Network Exchange (ONNX) · NVIDIA TensorRT-LLM · Inference Latency · Post-Training Quantization · NVIDIA JetPack SDK · Inference Deployment
Sources
NVIDIA TensorRT
As of
2025-09

See it in the full glossary →