Embodied AI Glossary中文

MASt3R

Advanced

Naver’s model that adds a matching head onto DUSt3R, outputting both a 3D point map and dense matching features.

MASt3R was proposed by Vincent Leroy, Yohann Cabon, and Jérôme Revaud at Naver in 2024, published at ECCV 2024. Its starting premise is that image matching — finding the pixels in two images that correspond to the same physical point — is fundamentally a 3D problem, tightly bound up with camera pose and scene geometry. It builds on DUSt3R, which takes two images and directly regresses a 3D coordinate for every pixel (a point map), robust to large viewpoint changes but with limited matching accuracy — and adds a new head that outputs dense local features, trained with an additional matching loss, paired with a fast mutual-nearest-neighbor matching algorithm that speeds up matching by several orders of magnitude. It substantially leads on several matching benchmarks, with a 30-point absolute improvement in VCRE AUC on the Map-free localization dataset. Both MASt3R-SfM and MASt3R-SLAM are built around it. The code is released under a CC BY-NC-SA 4.0 license, non-commercial use only.

ExampleGiven two tabletop photos of a robot’s workspace taken from very different viewpoints, MASt3R can directly output the pixel correspondences between the two images along with each one’s 3D point map, from which the relative pose between the two viewpoints can be computed.

Also called
MASt3R: Grounding Image Matching in 3D
Related
DUSt3R · MASt3R-SLAM · Feature Matching · Pointmap · Feed-Forward 3D Reconstruction · VGGT
Sources
arXiv: Grounding Image Matching in 3D with MASt3R
GitHub: naver/mast3r

See it in the full glossary →