Embodied AI Glossary中文

π³ (Pi3)

π³(Pi3)Advanced

A feed-forward 3D reconstruction model with no reference viewpoint, giving the same result no matter what order the images come in.

π³ is a feed-forward 3D reconstruction model released in July 2025 by the Shanghai AI Laboratory, Zhejiang University, and other institutions, with its repository marked ICLR 2026. Earlier methods such as DUSt3R and VGGT need one image designated as a reference viewpoint, with every result expressed in that image’s coordinate frame, so a poorly chosen reference image degrades the reconstruction. π³ instead uses a fully permutation-equivariant architecture (shuffle the input order and the output just follows the same reordering, unchanged in content), so it needs no reference frame at all, directly predicting an affine-invariant camera pose and a scale-invariant local pointmap for every image. The paper reports state-of-the-art results at the time on camera pose estimation, monocular and video depth estimation, and dense pointmap reconstruction. An upgraded version, Pi3X, released in December 2025, can take known poses, intrinsics, or depth as input and produce output at approximately true scale. The code is BSD-licensed; the weights are non-commercial only.

ExampleThe same set of 10 indoor photos is fed into π³ in different orders; the resulting point cloud and relative camera poses stay essentially the same, whereas a reference-frame-based method’s result changes depending on which image is chosen first.

Also called
Pi3, Pi3X, Permutation-Equivariant Visual Geometry Learning
Related
VGGT · DUSt3R · Feed-Forward 3D Reconstruction · Pointmap · Camera Extrinsics · Shanghai Artificial Intelligence Laboratory
Sources
π³: Permutation-Equivariant Visual Geometry Learning (arXiv)
yyfz/Pi3 (GitHub)
As of
2025-12

See it in the full glossary →