DUSt3R
AdvancedA feed-forward reconstruction model that regresses a per-pixel 3D point map directly from two images, with no camera parameters needed.
DUSt3R is a 3D reconstruction model proposed by Shuzhe Wang and colleagues at Naver Labs Europe and Aalto University, published at CVPR 2024. The traditional pipeline — structure from motion (SfM) plus multi-view stereo (MVS) — first estimates camera intrinsics and extrinsics, then triangulates matched points, a process with many steps that’s easy to break. DUSt3R flips this around: it feeds two images into a Transformer encoder-decoder and directly regresses a 3D coordinate for every pixel, called a point map, with both images’ points expressed in the first image’s camera coordinate frame and each carrying a confidence score; depth, pixel correspondence, relative pose, and focal length can all be read off the point map. With more than two images, a global alignment step then unifies all pairwise results into one coordinate system. DUSt3R is the landmark work in “feed-forward 3D reconstruction,” and later work like MASt3R, VGGT, and π³ all build on this approach. The code is released under a CC BY-NC-SA 4.0 non-commercial license.
ExampleSnapping two casual phone photos of a tabletop with no camera calibration, DUSt3R outputs the corresponding 3D point cloud for both images along with the two cameras’ relative pose, which can be used to quickly reconstruct a robot’s workspace.
- Also called
- DUSt3R: Geometric 3D Vision Made Easy, Dense and Unconstrained Stereo 3D Reconstruction
- Related
- Pointmap · MASt3R · VGGT · Feed-Forward 3D Reconstruction · Structure from Motion · Multi-View Stereo
- Sources
- arXiv 2312.14132: DUSt3R: Geometric 3D Vision Made Easy
naver/dust3r (GitHub) - As of
- 2024-06