SAM 3
SAM 3(可提示概念分割)CommonMeta's third-generation segmentation model that finds every instance matching a short phrase or an example image.
SAM 3 is Meta's third-generation segmentation model, released in November 2025, introducing the task of ‘promptable concept segmentation’: given a short noun phrase (such as ‘yellow school bus’), an example image patch, or both, the model finds every instance of that concept in an image or video, outputs a mask and identity for each one, and tracks them continuously. Before this, SAM and SAM 2 could only segment one user-selected target at a time and didn't accept text prompts. SAM 3 has about 848 million parameters and shares a vision encoder between a DETR-style detector and a memory-based video tracker, with a separate ‘presence head’ added to split judging ‘is it there’ from ‘where is it.’ It was released alongside the SA-Co benchmark, covering roughly 270,000 concepts. SAM 3.1, released in March 2026, sped up multi-object video tracking.
ExampleGiven a kitchen video and the text ‘cup,’ SAM 3 segments and continuously tracks every cup in frame, assigning each its own ID. Taking one cup's mask and combining it with the depth map then gives a point cloud for grasp-pose estimation.
- Also called
- Promptable Concept Segmentation, SAM 3.1
- Related
- Segment Anything Model · SAM 2 · Open-Vocabulary Segmentation · SAM 3D · Grounding DINO · Instance Segmentation
- Sources
- SAM 3: Segment Anything with Concepts (arXiv 2511.16719)
Meta AI Blog: Segment Anything Model 3
GitHub: facebookresearch/sam3 - As of
- 2026-03