Embodied AI Glossary中文

Metric Depth / Relative Depth

度量深度 / 相对深度Advanced

Metric depth gives true distance in meters; relative depth only gives near/far ordering, off by an unknown scale and offset.

Depth estimation output comes in two kinds. Metric depth (also called absolute depth) gives every pixel’s true distance from the camera, in meters — this is what a depth camera or lidar directly measures; relative depth only tells you what’s nearer and what’s farther, with values differing from true distance by an unknown scale and offset. A monocular image inherently has scale ambiguity: the same photo could be a nearby scale model or a distant real house, and different cameras have different focal lengths, so training on a mix of these without accounting for it creates contradictions. This is why methods like MiDaS use a loss insensitive to scale and offset, training on large mixed datasets to predict relative depth — generalizing well, but without metric scale — while ZoeDepth, Metric3D, and Depth Pro instead find ways to output metric depth directly. Robot grasping and obstacle avoidance need to know exactly how far an object is, so they need metric depth, or need a few real range measurements to align relative depth to true scale.

ExampleGiven the same tabletop photo, a relative-depth model can only tell you the cup is closer than the wall behind it; a metric-depth model gives roughly how many meters the cup is from the camera, which a robot arm needs in order to plan a grasp.

Also called
Absolute Depth, Scale Ambiguity, Affine-Invariant Depth
Related
Monocular Depth Estimation · Depth Estimation · Metric3D · Depth Anything · Depth Pro · Depth Camera
Sources
ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
MiDaS: Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer

See it in the full glossary →