DriveVLM
DriveVLM(快慢双系统智驾)AdvancedA 2024 Tsinghua and Li Auto autonomous-driving system where a vision-language model thinks slowly and a traditional planner executes fast.
DriveVLM is an autonomous-driving method proposed in February 2024 by Hang Zhao's group at Tsinghua University's Institute for Interdisciplinary Information Sciences together with Li Auto, published at CoRL 2024. Traditional autonomous-driving pipelines have limited ability to handle rare, complex long-tail scenarios. DriveVLM uses a vision-language model (VLM, built on Qwen-VL) to work through a chain of thought in three steps — scene description, scene analysis, and hierarchical planning — before finally outputting a driving trajectory. But a VLM isn't precise enough at spatial localization and is slow to run inference, so the paper also proposes DriveVLM-Dual: the VLM outputs a rough reference trajectory at low frequency (the slow system), while a traditional perception-and-planning module refines it into the actual trajectory at high frequency (the fast system), the two working together asynchronously. The paper reports the system has been deployed in production vehicles, with average inference of about 410 milliseconds on an onboard platform with two Orin X chips. This division of labor matches the fast-slow dual-system architecture also seen in robotics, such as Figure's Helix.
ExampleWhen it encounters an unusual road situation, the VLM first describes the scene, identifies key objects that could affect the vehicle and analyzes their impact, then gives a high-level decision such as “slow down and go around” along with a rough trajectory, which the traditional planner refines in real time into an executable trajectory.
- Also called
- DriveVLM-Dual, The Convergence of Autonomous Driving and Large Vision-Language Models
- Related
- Dual-System Architecture (System 1 / System 2) · Autonomous Driving · Vision-Language Model · Chain-of-Thought · Long-tail Problem · Qwen-VL
- Sources
- DriveVLM (arXiv 2402.12289)
DriveVLM 项目主页 (Chinese) - As of
- 2024-11