Embodied AI Glossary中文

TrackVLA

银河通用 TrackVLAAdvanced

A Galbot VLA for embodied visual tracking that simultaneously recognizes and plans a path to follow a target, using only its own first-person camera.

TrackVLA was released in May 2025 by Peking University's EPIC Lab (He Wang's group) together with Galbot and other institutions, published at CoRL 2025, a vision-language-action (VLA) model. The task is embodied visual tracking: using only its own first-person camera, the robot must recognize a specified target in a dynamic environment and keep following it. Earlier approaches split 'recognizing the target' and 'planning where to go' into two separate modules, which lets errors compound; TrackVLA instead has both share a single large-language-model backbone (Vicuna-7B), with a language head handling recognition and an anchor-based diffusion model outputting the movement trajectory. Training uses about 1.7 million samples, half of them tracking examples from the team's own EVT-Bench simulation benchmark and half video question-answering recognition samples. It has been deployed on a real Unitree Go2 quadruped, with the model running on a remote RTX 4090 server at roughly 10 frames per second, and it is reasonably robust to occlusion and fast target movement.

ExampleTold to 'follow the person in the blue shirt,' a Unitree Go2 running only a single head-mounted RGB camera recognizes the target in a crowd and keeps following it, and can pick the target back up after it is briefly blocked from view.

Also called
Galbot TrackVLA, TrackVLA: Embodied Visual Tracking in the Wild
Related
Embodied Visual Tracking · Vision-Language-Action Model · NavFoM (Galbot) · Navigation · Diffusion Model · Unitree Go2
Sources
TrackVLA: Embodied Visual Tracking in the Wild (arXiv 2505.23189)
TrackVLA project page
As of
2025-05

See it in the full glossary →