Prompt Depth Anything
AdvancedUsing a cheap lidar’s sparse depth readings as a “prompt” so a large depth model outputs accurate 4K metric depth.
Prompt Depth Anything was proposed by teams from Zhejiang University and ByteDance Seed, among others, published at CVPR 2025, with code and models released under the Apache-2.0 license. Monocular depth models such as Depth Anything capture fine edges well but don’t know the real-world scale; the low-cost lidar sensors built into devices like the iPhone give true distances but at very low resolution — 24×24 in the paper’s example. Prompt Depth Anything treats the lidar depth as a prompt and fuses it into Depth Anything at multiple scales inside the decoder, producing metric (real-world-scale) depth at up to 4K resolution. The paper includes grasping experiments on a Unitree H1: a policy trained only on diffuse (non-reflective, non-transparent) objects was able to grasp transparent and reflective objects once given this depth, outperforming versions that used only RGB or only lidar.
ExampleAn iPhone records RGB and lidar data at the same time; Prompt Depth Anything turns them into high-resolution depth, which is then converted into a point cloud for a grasping policy to use.
- Also called
- PromptDA, Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
- Related
- Depth Anything · Monocular Depth Estimation · Metric Depth / Relative Depth · Depth Completion · LiDAR · Transparent & Reflective Object Perception
- Sources
- Prompt Depth Anything 项目主页 (Chinese)
DepthAnything/PromptDA (GitHub)
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation (arXiv 2412.14015) - As of
- 2025-06