FoundationStereo
AdvancedNVIDIA’s stereo-matching foundation model that produces depth in a new scene with no fine-tuning needed.
FoundationStereo is a stereo-matching model proposed by Bowen Wen and colleagues at NVIDIA, published at CVPR 2025. Given left and right camera images, it outputs a dense disparity map, which is then converted to depth or a point cloud using the baseline and focal length. Earlier stereo-matching networks usually needed re-fine-tuning to work in a new scene; FoundationStereo is trained on about 1 million pairs of photorealistic synthetic stereo images (the FSD dataset) and uses side-tuning to bring in monocular-depth priors from a vision foundation model, narrowing the sim-to-real gap and achieving zero-shot generalization — NVIDIA said it ranked first on the Middlebury and ETH3D leaderboards at release. In robotics it’s commonly used to get more complete depth from stereo images for grasping and pose estimation; a real-time version, Fast-FoundationStereo, followed in December 2025.
ExamplePhotographing a tabletop with a stereo camera and feeding the left and right images into FoundationStereo produces a depth map and point cloud, which is then passed to FoundationPose to estimate a target object’s 6D pose.
- Also called
- FoundationStereo: Zero-Shot Stereo Matching, Fast-FoundationStereo
- Related
- Stereo Matching · Disparity · Stereo Camera · Depth Estimation · FoundationPose · Depth Anything
- Sources
- FoundationStereo: Zero-Shot Stereo Matching (arXiv 2501.09898)
NVlabs/FoundationStereo GitHub - As of
- 2025-12