LingBot-Depth
蚂蚁灵波 LingBot-DepthAdvancedAn open-source depth-completion model from Ant Group’s Robbyant that uses a color image to repair a depth camera’s holes and noise.
LingBot-Depth is a depth model released and open-sourced in January 2026 by Robbyant, the embodied-AI company under Ant Group; the paper is titled “Masked Depth Modeling for Spatial Perception,” and the GitHub page states it has been accepted at ECCV 2026. RGB-D depth cameras often can’t measure depth on transparent or reflective surfaces, leaving holes and noise. LingBot-Depth treats these missing regions as a natural “mask”: given a color image, raw depth, and camera intrinsics, it uses a ViT-Large backbone to fuse the two modalities and outputs a completed metric depth map and a point cloud in the camera’s coordinate frame. It was trained on about 3 million paired RGB-D samples (about 2 million real, 1 million synthetic). The team states its depth-completion error is 40–50% lower than the best prior methods. Code, weights, and the dataset are all open source.
ExampleIn the official grasping experiments, grasping hard-to-sense objects using depth repaired by LingBot-Depth: success rate for a transparent storage box rose from 0% to 50%, for a glass cup from 60% to 80%, and for a steel cup from 65% to 85%.
- Also called
- Masked Depth Modeling for Spatial Perception, Masked Depth Modeling
- Related
- Depth Completion · Depth Camera · Transparent & Reflective Object Perception · Depth Holes · Masked Autoencoder · Robbyant
- Sources
- Masked Depth Modeling for Spatial Perception (arXiv 2601.17895)
GitHub: Robbyant/lingbot-depth
Robbyant 官网:LingBot-Depth (Chinese) - As of
- 2026-09