TraceVLA
AdvancedA method that draws the recent motion trace of key points onto the image as a prompt, boosting a VLA's sense of space and time.
TraceVLA was released by the University of Maryland and Microsoft Research (Ruijie Zheng, Jianwei Yang, and others) in December 2024, published at ICLR 2025. VLAs usually see only a single current frame and have no sense of how the robot and objects have just been moving. TraceVLA introduces 'visual trace prompting': the point-tracking model CoTracker follows key points in the image over the past several frames, and their trace is drawn directly onto the image, so the model receives both the original frame and the frame with the trace overlaid. The authors used this to fine-tune OpenVLA on 150,000 manipulation trajectories they collected, producing TraceVLA, which beats OpenVLA by about 10 points across 137 configurations in SimplerEnv and reaches 3.5 times OpenVLA's performance on 4 real-robot WidowX tasks. A smaller version based on the 4-billion-parameter Phi-3-Vision comes close to the 7-billion-parameter OpenVLA.
ExampleA robot arm is moving a spoon toward a towel; the input image has a colored trace overlaid showing where the gripper and spoon were in the last few steps, letting the model judge the direction and progress of the motion when deciding the next action.
- Also called
- Visual Trace Prompting, TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
- Related
- OpenVLA · Visual Prompting · Tracking Any Point · CoTracker · SimplerEnv · Vision-Language-Action Model
- Sources
- TraceVLA (arXiv 2412.10345)
TraceVLA project page - As of
- 2025-01