NVIDIA TensorRT Edge-LLM
TensorRT Edge-LLMAdvancedNVIDIA's on-device inference framework for running large language and vision-language models on automotive and robotics chips.
TensorRT Edge-LLM is an open-source C++ inference framework from NVIDIA, built to run large language models and vision-language models on embedded platforms such as Jetson Thor and DRIVE AGX Thor. NVIDIA's data-center library for this, TensorRT-LLM, depends on a Python environment and targets high throughput across many GPUs; automotive and robotics deployments instead need low per-request latency, a small memory footprint, and lightweight deployment without Python, which is what Edge-LLM is built for. It reportedly supports low-precision quantization formats such as FP8 and NVFP4, along with speculative decoding (predicting several tokens ahead to speed up generation). For embodied AI, it is one option for running a VLA's or VLM's “brain” directly on the robot itself, rather than on a remote server.
ExampleQuantize a multi-billion-parameter VLM and deploy it on a Jetson Thor with TensorRT Edge-LLM, so the robot can do scene understanding and task planning locally.
- Also called
- Edge-LLM
- Related
- NVIDIA TensorRT · NVIDIA TensorRT-LLM · NVIDIA Jetson Thor · On-Device / Edge Deployment · On-device Model · Post-Training Quantization
- Sources
- NVIDIA/TensorRT-Edge-LLM (GitHub)
- As of
- 2026-01