Segment Anything Model (SAM v1)

← Back to Image Segmentation

SAM [19] is a large-scale segmentation foundation model trained by Meta AI on over 1 billion masks from 11 million images. It supports three prompting modes:

  • Point prompts: one or more foreground/background point clicks.

  • Box prompts: a bounding box around the target object.

  • Automatic mask generation: dense grid of point prompts to segment everything in the image.

SAM uses a large Vision Transformer (ViT) image encoder, a lightweight prompt encoder, and a fast mask decoder. The image embeddings are computed once per image and cached, enabling interactive prompting at near-real-time speeds.

See also

  • SAM 2 — SAM extended with a memory mechanism for video and improved mask quality

  • U-Net — a trained-from-scratch alternative when domain-specific labels are available

  • Image Segmentation — every other segmentation and tracking approach