U-Net¶
U-Net [17] is the dominant architecture for biomedical and scientific image segmentation. Its defining feature is a symmetric encoder–decoder structure with skip connections:
Encoder (contracting path): a stack of convolutional blocks with max-pooling that progressively reduces spatial resolution while increasing feature dimensionality.
Decoder (expanding path): transposed convolutions or bilinear upsampling restore spatial resolution, with skip connections concatenating high-resolution encoder features at each scale.
Output: a \(1 \times 1\) convolution producing per-pixel class logits.
The skip connections allow the decoder to recover fine spatial detail that would otherwise be lost in the bottleneck, which is critical for precise boundary localization.
The segmentation_models_pytorch library [18]
provides a high-level API for U-Net and many other architectures with
a choice of encoder backbones (ResNet, EfficientNet, etc.) pretrained
on ImageNet.
See also
Segment Anything Model (SAM v1) — a foundation-model alternative that needs no training
NASA MicroNet — domain-pretrained encoder backbones for U-Net-style models
Image Segmentation — every other segmentation and tracking approach