On-Device / Edge Deployment
端侧部署CommonRunning model inference directly on a robot's own onboard computer, instead of calling a remote server.
On-device deployment means model inference runs on the robot itself or on a device right next to it, typically an embedded board such as Jetson or RK3588. The alternative is streaming images to a cloud or lab GPU server and streaming actions back. Running on-device keeps latency low, works without a network connection, and keeps data from leaving the device — at the cost of limited compute, memory, and power, which often means large models need to be quantized, pruned, or distilled and then compiled with tools like TensorRT or RKNN before they can hit an adequate frame rate. Real systems often combine both approaches: low-level locomotion control on-device, large-model inference in the cloud or on an edge server, an arrangement known as cloud-edge-device collaboration.
ExampleGemini Robotics On-Device is a version of the VLA specifically optimized to run locally on a robot.
- Also called
- edge deployment, local deployment, onboard deployment
- Related
- On-device Model · Inference Latency · Cloud-Edge-Device Collaboration · Policy Server (Remote Inference) · Post-Training Quantization · NVIDIA Jetson
- Sources
- Edge computing - Wikipedia