Embodied AI Glossary中文

MapAnything

Advanced

Meta and CMU’s general-purpose feed-forward 3D reconstruction model that directly outputs true-scale scene structure.

MapAnything was released by Meta Reality Labs and Carnegie Mellon University in September 2025, a Transformer-based feed-forward 3D reconstruction model: given one or more images, optionally supplemented with camera intrinsics, poses, depth, or partial reconstruction results, a single forward pass outputs each image’s depth map, local ray map, camera pose, and a unified metric scale factor — together forming a 3D scene at true, meter-based scale. Tasks that used to each need their own algorithm — uncalibrated structure from motion, multi-view stereo, monocular depth estimation, camera localization, depth completion — are all covered by this one model. It belongs to the same feed-forward reconstruction lineage as DUSt3R and VGGT, with code and weights open-sourced (including an Apache-licensed version). For robotics, it can turn a handful of photos into a metric point cloud that is directly usable for planning.

ExampleTaking a few phone photos around a tabletop with no camera parameters provided, MapAnything can output a metric-scale point cloud and each photo’s camera pose; if camera intrinsics are known, they can also be supplied as a constraint.

Also called
MapAnything: Universal Feed-Forward Metric 3D Reconstruction, Map Anything
Related
Feed-Forward 3D Reconstruction · DUSt3R · VGGT · MASt3R · Structure from Motion · Metric Depth / Relative Depth
Sources
arXiv: MapAnything: Universal Feed-Forward Metric 3D Reconstruction
MapAnything 项目主页 (Chinese)
As of
2026-01

See it in the full glossary →