Embodied AI Glossary中文

SAM 3

SAM 3(可提示概念分割)Common

Meta's third-generation segmentation model that finds every instance matching a short phrase or an example image.

SAM 3 is Meta's third-generation segmentation model, released in November 2025, introducing the task of ‘promptable concept segmentation’: given a short noun phrase (such as ‘yellow school bus’), an example image patch, or both, the model finds every instance of that concept in an image or video, outputs a mask and identity for each one, and tracks them continuously. Before this, SAM and SAM 2 could only segment one user-selected target at a time and didn't accept text prompts. SAM 3 has about 848 million parameters and shares a vision encoder between a DETR-style detector and a memory-based video tracker, with a separate ‘presence head’ added to split judging ‘is it there’ from ‘where is it.’ It was released alongside the SA-Co benchmark, covering roughly 270,000 concepts. SAM 3.1, released in March 2026, sped up multi-object video tracking.

ExampleGiven a kitchen video and the text ‘cup,’ SAM 3 segments and continuously tracks every cup in frame, assigning each its own ID. Taking one cup's mask and combining it with the depth map then gives a point cloud for grasp-pose estimation.

Also called
Promptable Concept Segmentation, SAM 3.1
Related
Segment Anything Model · SAM 2 · Open-Vocabulary Segmentation · SAM 3D · Grounding DINO · Instance Segmentation
Sources
SAM 3: Segment Anything with Concepts (arXiv 2511.16719)
Meta AI Blog: Segment Anything Model 3
GitHub: facebookresearch/sam3
As of
2026-03

See it in the full glossary →