Metric3D
AdvancedA model that estimates metric depth from a single image, resolving scale ambiguity across cameras with a canonical camera-space transform.
Metric3D was proposed by Wei Yin, Chunhua Shen, and colleagues, published at ICCV 2023, aiming to estimate depth at true scale from a single photo in a zero-shot setting. The difficulty is that different cameras have different focal lengths, so objects of the same size appear at different apparent distances on the image, and training on data mixed together from many cameras creates conflicting scale signals. Its solution is a “canonical camera-space transform”: during training, every sample is converted according to its focal length into a shared virtual canonical camera, and at inference time the result is converted back using the real camera’s intrinsics. This lets it train on over 8 million images from more than a thousand different cameras, giving metric depth even for unseen cameras, and it achieved the best results at the time on seven zero-shot benchmarks. The 2024 Metric3D v2 (TPAMI) switched to a DINOv2 backbone, expanded training data to over 16 million images, and added surface normal estimation. The code is open source under a BSD license.
ExampleMonocular SLAM using just an ordinary camera has no way to know the scene’s true scale; the Metric3D repository recommends an open-source implementation that feeds its predicted depth into DROID-SLAM to get a mapping result at metric scale.
- Also called
- Metric3D v2, Metric3Dv2
- Related
- Monocular Depth Estimation · Metric Depth / Relative Depth · Camera Intrinsics · Depth Anything · Depth Pro · MoGe
- Sources
- arXiv: Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image
arXiv: Metric3Dv2: A Versatile Monocular Geometric Foundation Model
GitHub: YvanYin/Metric3D - As of
- 2025-01