Embodied AI Glossary中文

CUT3R

Advanced

A 3D reconstruction model with a persistent memory state that reads images one at a time and updates the whole scene online.

CUT3R stands for Continuous Updating Transformer for 3D Reconstruction, proposed by researchers at UC Berkeley and Google DeepMind, and presented as an oral paper at CVPR 2025. It follows the same idea as DUSt3R-style models — regressing point maps directly from images, with one 3D point per pixel — but adds a continuously updated state: as each new image arrives, the model first uses it to update the state, then outputs that image’s point map in a shared coordinate system, at real-world scale; the reconstruction gradually gets more complete as more input arrives, with no per-video optimization needed. It can process either a video stream or an unordered set of photos, supports dynamic scenes with moving objects, and can even infer unseen regions from virtual viewpoints that were never actually photographed.

ExampleFeeding a robot head camera’s video into CUT3R frame by frame produces, for each new frame, a point map aligned into the same coordinate system; accumulated together, this becomes a room point cloud that keeps getting more complete.

Also called
Continuous 3D Perception Model with Persistent State, Continuous Updating Transformer for 3D Reconstruction
Related
DUSt3R · Pointmap · Feed-Forward 3D Reconstruction · 4D Reconstruction · VGGT · MASt3R-SLAM
Sources
Continuous 3D Perception Model with Persistent State (arXiv 2501.12387)
CUT3R project page
As of
2025-06

See it in the full glossary →