Segment Anything Model (SAM v1)¶
SAM [19] is a large-scale segmentation foundation model trained by Meta AI on over 1 billion masks from 11 million images. It supports three prompting modes:
Point prompts: one or more foreground/background point clicks.
Box prompts: a bounding box around the target object.
Automatic mask generation: dense grid of point prompts to segment everything in the image.
SAM uses a large Vision Transformer (ViT) image encoder, a lightweight prompt encoder, and a fast mask decoder. The image embeddings are computed once per image and cached, enabling interactive prompting at near-real-time speeds.
See also
SAM 2 — SAM extended with a memory mechanism for video and improved mask quality
U-Net — a trained-from-scratch alternative when domain-specific labels are available
Image Segmentation — every other segmentation and tracking approach