Embodied AI Glossary中文

VGGT

Advanced

A 3D vision model that computes camera parameters, depth, and a point cloud from multiple images in a single forward pass.

VGGT is a feed-forward 3D reconstruction model from Oxford’s Visual Geometry Group (VGG) and Meta AI, with about 1 billion parameters, and won the CVPR 2025 Best Paper Award. Given anywhere from one image to a few hundred, it runs a single forward pass to output every image’s camera parameters, depth map, pointmap (the 3D coordinate corresponding to each pixel), and 3D point tracks — usually in under a second, with no need for bundle adjustment (the repeated optimization of cameras and 3D points that traditional reconstruction relies on). Work that previously required a multi-stage pipeline such as COLMAP is compressed into a single network. In embodied AI, it is commonly used to quickly get scene geometry from multi-view images. The original weights are non-commercial only; in July 2025, a separate set of commercially usable weights and open training code were released.

ExampleTwenty phone photos taken circling a tabletop are fed into VGGT, producing each photo’s camera pose and a dense point cloud of the whole table in about a second.

Also called
Visual Geometry Grounded Transformer, VGGT-1B
Related
Feed-Forward 3D Reconstruction · DUSt3R · π³ (Pi3) · Structure from Motion · Pointmap · Bundle Adjustment
Sources
VGGT: Visual Geometry Grounded Transformer (arXiv)
facebookresearch/vggt (GitHub)
As of
2025-07

See it in the full glossary →