MoGe
AdvancedA Microsoft Research model that estimates full 3D geometry from a single image, predicting a 3D point for every pixel.
MoGe is a monocular geometry estimation model from Microsoft Research, presented as an oral paper at CVPR 2025. From a single ordinary photo, it directly predicts a point map — a 3D coordinate for every pixel — along with a depth map and the camera’s field of view. The first version predicts an affine-invariant point map, correct only up to an unknown scale and translation, trained with a specially designed global-alignment loss plus a multi-scale local-geometry loss. MoGe-2, released in June 2025, predicts point maps at true metric (real-world) scale and adds surface-normal estimation. Robots can use it to recover a scene’s geometry from a single camera image, without needing multiple viewpoints or stereo.
ExampleGiven a photo of a kitchen found online, MoGe outputs a 3D coordinate for every pixel plus the camera’s field of view, which can be converted directly into a point cloud.
- Also called
- Monocular Geometry Estimation, MoGe-2
- Related
- Monocular Depth Estimation · Pointmap · Metric Depth / Relative Depth · Depth Anything · DUSt3R · Depth Pro
- Sources
- microsoft/MoGe - GitHub
CVPR 2025 Oral: MoGe - As of
- 2025-06