Embodied AI Glossary中文

Depth Anything 3

DA3Advanced

ByteDance Seed’s geometry model that recovers depth and camera poses from any number of images, with poses optional.

Depth Anything 3 was released by ByteDance’s Seed team in November 2025, the third generation of the Depth Anything series. The first two generations only did depth estimation from a single image; DA3 extends this to any number of input images, with camera poses either known or unknown, outputting spatially consistent depth and camera parameters, which can be further turned into a point cloud or 3D Gaussians. The design is deliberately simple: the backbone is just an ordinary Transformer (the original DINO encoder), and the training objective is unified into a single “depth plus ray” prediction. The paper reports, on its own visual geometry benchmark, that camera pose accuracy is on average 44.3% higher than the previous best method, VGGT, and geometric accuracy is 25.1% higher. Models range in size from Small (0.08B) to Giant (1.15B), with additional metric-depth and monocular-specific versions available.

ExampleWalking around a table taking 5 photos on a phone with no camera parameters provided, DA3 gives every image’s depth and camera pose in a single forward pass and fuses them into one tabletop point cloud.

Also called
DA3, Depth Anything 3: Recovering the Visual Space from Any Views
Related
Depth Anything · VGGT · Feed-Forward 3D Reconstruction · Monocular Depth Estimation · 3D Gaussian Splatting · DINOv2
Sources
Depth Anything 3: Recovering the Visual Space from Any Views (arXiv 2511.10647)
ByteDance-Seed/Depth-Anything-3 (GitHub)
As of
2025-12

See it in the full glossary →