Embodied AI Glossary中文

Structure from Motion

运动恢复结构SfMAdvanced

Computing both camera poses and a scene’s 3D points at once from a set of photos taken from different angles.

Structure from motion is a classic 3D reconstruction method: given photos of the same scene taken from multiple viewpoints (which can be in any order), it outputs the camera pose (position and orientation) for every photo along with a sparse 3D point cloud. The standard pipeline first extracts feature points and matches them across images, uses epipolar geometry to reject bad matches, then adds images one at a time while triangulating 3D points, and finally runs bundle adjustment — jointly fine-tuning all poses and 3D points to minimize reprojection error — as a global refinement. It differs from SLAM in that it usually runs offline and doesn’t require the images to be in chronological order. The open-source tool COLMAP is the most commonly used implementation; the first step in building a NeRF, a 3D Gaussian Splat, or bringing a real scene into simulation (real-to-sim) is often to compute camera poses with SfM. Recently, feed-forward models such as DUSt3R and VGGT have tried to output both pose and geometry directly, in one shot.

ExampleFifty phone photos taken while circling a mug on a table are fed into COLMAP, producing each photo’s camera pose and a sparse point cloud of the mug, which is then used to train a 3D Gaussian Splat.

Also called
SfM
Related
Bundle Adjustment · Feature Matching · Triangulation · COLMAP · Multi-View Stereo · Feed-Forward 3D Reconstruction
Sources
Structure-from-Motion Revisited (Schönberger & Frahm, CVPR 2016)
COLMAP 官方文档 (Chinese)

See it in the full glossary →