VGGT
AdvancedA 3D vision model that computes camera parameters, depth, and a point cloud from multiple images in a single forward pass.
VGGT is a feed-forward 3D reconstruction model from Oxford’s Visual Geometry Group (VGG) and Meta AI, with about 1 billion parameters, and won the CVPR 2025 Best Paper Award. Given anywhere from one image to a few hundred, it runs a single forward pass to output every image’s camera parameters, depth map, pointmap (the 3D coordinate corresponding to each pixel), and 3D point tracks — usually in under a second, with no need for bundle adjustment (the repeated optimization of cameras and 3D points that traditional reconstruction relies on). Work that previously required a multi-stage pipeline such as COLMAP is compressed into a single network. In embodied AI, it is commonly used to quickly get scene geometry from multi-view images. The original weights are non-commercial only; in July 2025, a separate set of commercially usable weights and open training code were released.
ExampleTwenty phone photos taken circling a tabletop are fed into VGGT, producing each photo’s camera pose and a dense point cloud of the whole table in about a second.
- Also called
- Visual Geometry Grounded Transformer, VGGT-1B
- Related
- Feed-Forward 3D Reconstruction · DUSt3R · π³ (Pi3) · Structure from Motion · Pointmap · Bundle Adjustment
- Sources
- VGGT: Visual Geometry Grounded Transformer (arXiv)
facebookresearch/vggt (GitHub) - As of
- 2025-07