Multi-View Stereo
多视图立体MVSAdvancedA method that recovers a scene’s dense 3D structure from multiple photos whose camera positions are already known.
Multi-view stereo (MVS) is a classic 3D reconstruction technique. Given each photo’s camera intrinsics and extrinsics — usually solved first by structure from motion (SfM) — it finds correspondences for the same physical point across multiple images, then triangulates to compute depth, producing a dense depth map or point cloud that can be turned into a mesh. SfM alone yields only a sparse set of feature points; MVS is what fills that in to a dense model. COLMAP is a widely used implementation of the classic pipeline; since 2018, deep-learning methods such as MVSNet have used neural networks to regress depth directly. MVS is a basic building block for digital twins, scanning real objects into simulation assets, and novel view synthesis.
ExampleCOLMAP first runs structure from motion on a set of photos taken around an object to recover the camera poses, then runs multi-view stereo to produce a dense point cloud.
- Also called
- MVS, Multi-View 3D Reconstruction
- Related
- Structure from Motion · COLMAP · Triangulation · Bundle Adjustment · Neural Radiance Fields · 3D Gaussian Splatting
- Sources
- COLMAP Tutorial
MVSNet: Depth Inference for Unstructured Multi-view Stereo (arXiv)