Embodied AI Glossary中文

ScanNet

ScanNet 数据集Advanced

A large-scale indoor RGB-D scan dataset with 3D reconstruction and semantic annotations.

ScanNet is an indoor-scene dataset released by Angela Dai, Matthias Nießner, and colleagues at CVPR 2017: 1,513 indoor scenes were scanned with consumer RGB-D cameras (capturing color and depth simultaneously), totaling 2.5 million frames, along with camera poses, 3D surface reconstructions, and crowdsourced semantic segmentation labels, providing a unified training and evaluation set for 3D semantic segmentation and 3D detection. In embodied AI, many 3D-language tasks are built on top of it, such as 3D visual grounding (locating an object in a 3D scene from a sentence) and 3D question answering. A later version, ScanNet++, switched to a laser scanner, a DSLR, and an iPhone for capture, achieving higher precision, and v2, released in December 2024, expanded coverage to more than 1,000 scenes.

ExampleScanRefer writes 51,583 natural-language descriptions for 11,046 objects across 800 ScanNet scenes; a model must use the description to locate the target object within the 3D scan.

Also called
ScanNet++, ScanNet v2
Related
3D Visual Grounding · Depth Camera · Point Cloud · 3D Scene Graph · Habitat-Matterport 3D Dataset · Matterport3D
Sources
ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes (arXiv 1702.04405)
ScanNet++ 官网 (Chinese)
ScanRefer (arXiv 1912.08830)
As of
2024-12

See it in the full glossary →