Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

3.10 Instance Segmentation with YOLO

Open In Colab

Overview

In this lesson, we will train an Instance Segmentation model using YOLO26. You’ll be able to choose specific augmentations, batch size, resolution, and other parameters based on your system’s capabilities and runtime. The dataset is already provided in YOLO format and will be used to train and evaluate the model.

Learning Objectives

By the end of this section, you will:

  • Understand the extended YOLO format and how to train a custom instance segmentation model using YOLO26.

  • Experiment with different augmentations and hyperparameters for instance segmentation.

  • Evaluate the model’s performance and visualize bounding boxes and masks.

Background

For instance segmentation in YOLO, each object is still labeled with a bounding box (as in standard object detection) and an associated segmentation mask. In practice, the YOLO label files include extra information to store the polygonal representation or encoded mask for each object. This is unique compared to many other frameworks because YOLO’s model architecture can simultaneously predict bounding boxes for localization and generate masks for precise instance segmentation.

Downloading the Dataset

The dataset for this lesson is already formatted in the YOLO instance segmentation format. You can load it directly for training and evaluation. Ensure you have the dataset uploaded before proceeding.

Access the dataset at the following link:

https://universe.roboflow.com/lini-foundation/lini-coral-forms-1.0/dataset/1

Preparing the Environment

Let’s first install the required libraries and set up the environment to train our YOLO26 instance segmentation model. Crucially, make sure that you are in a GPU runtime by running the cell below. It should output the GPU currently connected to.

Dataset Format: YOLO Segmentation Annotations

Before unzipping, confirm your dataset is structured as a zip file containing image files alongside YOLO-format segmentation label files (one .txt per image). YOLO segmentation format differs from the detection format you may have used in earlier lessons. A detection label line looks like this:

class_id cx cy w h

where cx cy is the normalized bounding-box center and w h is its normalized width and height. A segmentation label line replaces the box with a polygon:

class_id px1 py1 px2 py2 ... pxN pyN

Here each (px, py) pair is a normalized coordinate (values between 0 and 1) for one vertex of the polygon tracing the object outline. The number of vertex pairs varies by object: a roughly rectangular fish might need 6 to 8 pairs, while a complex coral structure could require many more. Tools such as Label Studio and Roboflow can export annotations directly in this format.

Initializing TensorBoard Before Training

TensorBoard is a powerful visualization tool that provides real-time insights into your model’s training process. By initializing TensorBoard before training, you can monitor key metrics such as loss, accuracy, and learning rates, allowing for timely adjustments and improved model performance. This proactive monitoring helps in identifying issues like overfitting or underfitting early in the training process. Be sure to click the refresh button in the top right of the Tensorboard often!

Training the Instance Segmentation Model

This cell loads the pre-trained yolo26n-seg.pt weights and fine-tunes them on your coral dataset for 100 epochs. The n in the model name stands for “nano,” the smallest and fastest variant in the YOLO26 family. That makes it a reasonable starting point: training finishes quickly enough to give you meaningful feedback on whether your dataset and annotation quality are sound before you commit to longer runs with larger variants such as yolo26s-seg.pt or yolo26m-seg.pt.

The data argument points to the data.yaml file extracted from your zip. That file tells the trainer where to find training and validation images, how many classes you have, and what each class is called. If training raises a FileNotFoundError, the most common cause is a path mismatch between what data.yaml specifies and where the dataset actually landed after unzipping: open the file and verify the train, val, and test paths before re-running.

Training outputs accumulate under /content/runs/segment/train/. The best checkpoint, saved by validation loss, will be written to weights/best.pt. TensorBoard (launched in the cell above) reads from that same directory, so metrics should appear there as soon as the first epoch completes.

Visualizing Predictions on Test Images

After training finishes, this cell loads the best checkpoint and runs inference on every image in the test split. For each image, it does two things. First, it calls the YOLO model and saves annotated output images (with masks and bounding boxes drawn) to /content/runs/segment/predict. Second, it computes a simple coverage statistic: all instance masks are merged into a single binary mask, and the fraction of the image that mask covers is reported as a percentage.

That coverage number is not a biological measurement on its own, but it is a useful sanity check. If you are segmenting a reef transect image and the model reports 0.3% coverage, something has almost certainly gone wrong (either no detections or very tight masks). If it reports 85% coverage on an image that is mostly open water, the masks are probably overflowing their targets. Use the Plotly figure titles to scan quickly before doing any deeper analysis.

The px.imshow call renders each annotated image as an interactive Plotly figure. You can zoom in to inspect mask boundaries, which is especially useful for evaluating whether the model cleanly separates adjacent coral colonies or fish that are partially occluding one another.


image 1/14 /content/dataset/test/images/20230704_174611_mp4-8_jpg.rf.e616459a2ba009da3e97af91ee6896e9.jpg: 640x640 1 phaceolid, 12.0ms
image 2/14 /content/dataset/test/images/20230829_123255_jpg.rf.83ef6c419fb9450cc30533be3a0f3152.jpg: 640x640 1 Parascolymia, 10.8ms
image 3/14 /content/dataset/test/images/20230829_123544_jpg.rf.68e223d29096ccf784464c9d4e08160f.jpg: 640x640 1 Parascolymia, 11.1ms
image 4/14 /content/dataset/test/images/20230829_123650_jpg.rf.4e4157deb8ec4d1e4574e9414e8f10ce.jpg: 640x640 1 Parascolymia, 10.8ms
image 5/14 /content/dataset/test/images/20230829_124126_jpg.rf.8e7d988292723a46eb3283dca055a962.jpg: 640x640 1 phaceolid, 10.8ms
image 6/14 /content/dataset/test/images/20230829_124409_jpg.rf.871e34845775219697ebe21e97cb13c3.jpg: 640x640 1 phaceolid, 10.7ms
image 7/14 /content/dataset/test/images/20230829_124510_jpg.rf.10f413fa7f82a5196905071454271061.jpg: 640x640 1 phaceolid, 11.0ms
image 8/14 /content/dataset/test/images/20230829_124619_jpg.rf.f539c120bb7a0c8568a490654b824394.jpg: 640x640 1 massive, 11.5ms
image 9/14 /content/dataset/test/images/20230829_124712_jpg.rf.cbcb6be5d2e9999b1aad9cd87fc43f11.jpg: 640x640 1 massive, 11.9ms
image 10/14 /content/dataset/test/images/20230829_124915_jpg.rf.7df15bfc9f192d8bf1811121154b0871.jpg: 640x640 1 submassive, 11.4ms
image 11/14 /content/dataset/test/images/20230829_125036_jpg.rf.fc23a71e88b2f989140870a79a9c2ea3.jpg: 640x640 1 Encrusting, 1 branching, 11.1ms
image 12/14 /content/dataset/test/images/20230829_125447_jpg.rf.7f40cac22f123bca6e074adc2c4f7a19.jpg: 640x640 1 branching, 11.2ms
image 13/14 /content/dataset/test/images/20230829_130442_jpg.rf.e8a337d3a2f754e0862ac1dd807302c4.jpg: 640x640 1 submassive, 11.0ms
image 14/14 /content/dataset/test/images/20230829_130539_jpg.rf.1b6b1a3dceee3da29739ade4bcca6969.jpg: 640x640 2 branchings, 11.2ms
Speed: 2.3ms preprocess, 11.2ms inference, 1.9ms postprocess per image at shape (1, 3, 640, 640)
Results saved to runs/segment/predict3
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...

Reflecting on Results

Instance segmentation produces a polygon mask for each detected object rather than a simple rectangular boundary. This extra precision matters for downstream analysis: you can count individuals, measure each organism’s pixel area, and track specific animals across frames without conflating overlapping neighbors. Bounding boxes tell you where something is; masks tell you exactly which pixels belong to it, which is the foundation for any morphometric or density estimate you want to extract automatically.

Spend some time reviewing the outputs from the final visualization cell. Check whether the masks follow the actual body outline of each organism or whether they collapse into rough approximations that miss fins, tails, or protruding features. Pay particular attention to crowded frames where two animals overlap: note whether the model correctly separates their masks or whether one mask bleeds into the adjacent object.

Reflection questions:

  1. How does an instance segmentation mask differ from a simple bounding box in terms of information content, and what analyses become possible once you have per-pixel labels?

  2. For your marine dataset, which species or object types are most difficult to segment cleanly, and why (consider body shape, coloration, motion blur, or density of individuals)?

  3. How might you use mask area (measured in pixels and converted to real-world units via camera calibration) as a proxy for biomass or organism size, and what sources of error would you need to account for?