What is semantic segmentation
Semantic segmentation assigns a class label to every pixel in an image. While classification gives one label to the whole image and object detection returns bounding boxes, semantic segmentation outputs a per-pixel “class map” with the same spatial size as the input.
How it differs from classification and detection
Classification answers “what is in the image”; detection answers “which objects are present and where” (with boxes); semantic segmentation answers “which class each pixel belongs to”. It produces finer boundaries, but it does not separate different instances of the same class (that is instance segmentation).
From fully convolutional networks to encoder-decoder
In 2015 Long et al. proposed the Fully Convolutional Network (FCN), replacing the classifier’s fully connected layers with convolutions so the network can take arbitrary-size inputs and output a full-resolution segmentation map. U-Net (Ronneberger et al., 2015) later used a symmetric encoder-decoder with skip connections to combine shallow detail with deep semantics. The DeepLab family introduced dilated convolutions to enlarge the receptive field and refined edges with CRFs.
Key building blocks
Typical components include downsampling via pooling or strided convolution, upsampling via transposed convolution or bilinear interpolation, skip connections to recover spatial detail, dilated convolutions for a larger receptive field, and attention for global context. After Transformers, models such as SETR and SegFormer brought self-attention to segmentation.
How it is evaluated
The most common metric is IoU (intersection over union), and the mean IoU (mIoU) across classes is the standard benchmark. Pixel accuracy and the Dice coefficient are also used.
Applications
Road, lane and drivable-area segmentation in autonomous driving; organ and lesion delineation in medical imaging; land-cover mapping in remote sensing; defect localization in industrial inspection; and matting or background replacement in image editing.
Limitations
Pixel-level annotation is expensive; small objects, occlusion, class imbalance and blurry edges hurt accuracy; and models often degrade under domain shift. Weakly-, semi- and self-supervised segmentation are active research directions.
Sources: compiled from publicly available academic papers and industry materials, including FCN (Long et al., 2015), U-Net (Ronneberger et al., 2015), the DeepLab papers, and public technical reviews.