Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

3.9 Keypoint Detection With YOLO

Open In Colab

Overview

Bounding boxes tell you where a fish is. Keypoint detection tells you something about its shape. In this lesson you will train a YOLO26 pose estimation model to locate two anatomical landmarks on black cod (sablefish, Anoplopoma fimbria): the snout tip and the tail tip. Once you have those two points, measuring body length in the image becomes a single Euclidean distance calculation. With a known scale factor, that pixel distance converts directly to centimeters.

This approach is a form of morphometrics, the quantitative study of body shape and size. Field biologists have traditionally measured fish by hand or from stereo-video rigs with calibrated baselines. A trained keypoint model can automate that pipeline for ordinary monocular camera footage, making large-scale length-frequency surveys practical without diver intervention.

What YOLO26 Pose Estimation Outputs

For each detected object the model returns one bounding box and a fixed set of keypoints. Every keypoint is a three-element tuple (x, y, v) where x and y are pixel coordinates (normalized to [0, 1] in the label format, absolute in inference output) and v is a visibility flag. The visibility convention used by YOLO follows COCO: 0 means the keypoint was not labeled, 1 means it was labeled but may be occluded, and 2 means it is labeled and fully visible.

Learning Objectives

By the end of this lesson you will be able to:

  • Describe the YOLO keypoint annotation format and how visibility flags work.

  • Prepare a labeled dataset for a two-keypoint fish pose model.

  • Train yolo26n-pose.pt on a custom dataset and interpret training metrics.

  • Run inference and extract keypoint coordinates from the results object.

  • Compute pixel-space body length from keypoint pairs and discuss how to calibrate that to real-world units.

this is a test

Environment Setup

This notebook runs on Google Colab with a GPU runtime. Before doing anything else, go to Runtime > Change runtime type and select a GPU accelerator (T4 is sufficient for the nano model). Training on CPU is possible but will be many times slower and is not recommended.

The only external dependency beyond standard Colab packages is ultralytics, which bundles YOLO26 along with its training loop, data loaders, and inference utilities. The cell below confirms your GPU is visible to the system, followed by the package install.

Dataset and Annotation Format

YOLO Keypoint Label Format

YOLO keypoint labels are plain text files, one per image, with each line describing one object instance. The format is:

class_id  cx  cy  w  h  kp1x  kp1y  kp1v  kp2x  kp2y  kp2v  ...
  • class_id: integer class index (0 for black cod if it is your only class).

  • cx cy w h: bounding box center coordinates and dimensions, all normalized to [0, 1] relative to image width and height.

  • kp1x kp1y kp1v: x coordinate, y coordinate, and visibility flag for the first keypoint, also normalized.

  • Additional keypoint triplets follow in the same order for every instance.

The visibility flag v uses COCO conventions: 0 means the keypoint was not annotated at all, 1 means it was annotated but might be occluded or off-frame, and 2 means it is annotated and clearly visible. During training the model learns to suppress confidence for keypoints flagged as 0, so accurate flagging during annotation directly affects what the model learns.

Keypoints for Black Cod Length Measurement

For this lesson, each fish instance has exactly two keypoints:

  1. Keypoint 0: Snout tip (the anteriormost point of the upper jaw).

  2. Keypoint 1: Tail tip (the posteriormost point of the caudal fin, taken at the midline).

The straight-line pixel distance between these two points approximates standard length (snout to caudal peduncle) only if the fish is lying flat relative to the camera. For fish at an angle to the camera plane, the measured distance will be a foreshortened projection of true length, which is an important caveat when interpreting results.

The data.yaml file for a two-keypoint pose dataset includes a kpt_shape field that tells the model how many keypoints to expect and whether visibility flags are included:

kpt_shape: [2, 3]   # 2 keypoints, each with x, y, visibility

Obtaining Labeled Data

You can annotate images using Roboflow, which supports keypoint export directly in YOLO pose format. Alternatively, label with a tool such as CVAT and export to COCO keypoints JSON, then convert. The next cell shows the Roboflow download pattern. Replace the stub with your own project credentials or upload a pre-labeled dataset zip to /content/.

Training the Keypoint Model

Model Selection

YOLO26 ships a family of pose models scaled by parameter count: yolo26n-pose.pt (nano), yolo26s-pose.pt (small), yolo26m-pose.pt (medium), and so on. For an initial experiment on a modest dataset, the nano model trains quickly and still achieves reasonable keypoint localization accuracy. You can upgrade to a larger model once you have confirmed your data pipeline is correct and your annotations are consistent.

The model is initialized from COCO pose weights. COCO pose targets 17 human body keypoints, so the pretrained feature extractor will not have seen fish anatomy before. Fine-tuning replaces the detection and keypoint heads entirely, but the backbone weights give the model a strong starting point for recognizing edges, textures, and spatial relationships, which transfers reasonably well across domains.

Training Parameters

Key parameters to be aware of:

  • data: path to your data.yaml file. This file must include the kpt_shape: [2, 3] field for a two-keypoint pose task.

  • epochs: 100 is a reasonable starting point. Monitor the pose/loss curve in TensorBoard or the printed logs. If it has not plateaued by epoch 100, extend training.

  • imgsz: 640 pixels on the longer side. Underwater footage often has motion blur and low contrast, so do not reduce this below 480 without testing accuracy first.

  • batch: 16 fits comfortably on a T4 GPU. Increase to 32 on an A100 or reduce to 8 if you see out-of-memory errors.

Training artifacts, including weights and metric logs, are saved under /content/runs/pose/train/ by default.

Running Inference and Measuring Body Dimensions

Once training is complete, the best model checkpoint lives at /content/runs/pose/train/weights/best.pt. You load it with YOLO() just as you would any other YOLO model; the pose variant simply returns an additional .keypoints attribute on each result.

Extracting Keypoint Coordinates

The results object returned by inference is a list, one element per image. Each element carries:

  • r.boxes: bounding boxes, confidence scores, and class IDs.

  • r.keypoints: a Keypoints object. Calling .xy gives you a tensor of shape (num_detections, num_keypoints, 2) in absolute pixel coordinates. Calling .xyn gives normalized coordinates instead. Calling .conf gives per-keypoint confidence scores if available.

After moving the tensor to CPU and converting to a NumPy array with .cpu().numpy(), you can index directly: fish_kps[0] is the snout tip (x, y) and fish_kps[1] is the tail tip (x, y).

Computing Body Length

The Euclidean distance between snout and tail in pixel space is:

length_px = sqrt((tail_x - snout_x)^2 + (tail_y - snout_y)^2)

np.linalg.norm(fish_kps[0] - fish_kps[1]) computes exactly that. To convert to real-world units you need a scale factor in cm/pixel. Common approaches include:

  • Known object in frame: a ruler, calibration target, or laser scalers (two parallel dots at a fixed separation) placed in the same plane as the fish.

  • Stereo baseline: if two calibrated cameras capture the scene, depth from disparity gives you the fish-to-camera distance, from which you can derive cm/pixel using the camera’s focal length.

  • Fleet average: for comparative studies where absolute size matters less than relative size, you can normalize lengths to a reference individual or report everything in pixels with confidence intervals.

The inference cell below prints pixel-space body length. If you have a scale factor, multiply dist by that factor to obtain centimeters.

Reflecting on Results

Work through these questions after you have run inference and looked at the annotated output images saved under /content/runs/pose/predict/.

1. Calibration to real-world length. The model outputs pixel distances, but fishery science requires measurements in centimeters or millimeters. Describe a practical calibration method you could use with the camera system available to you. What information do you need to collect in the field, and how would you propagate measurement uncertainty through the calibration step to report a final length estimate with a confidence interval?

2. Additional keypoints for richer morphometrics. Snout-to-tail distance approximates standard length, but fish biologists often record several other measurements: fork length, body depth at the pectoral fin insertion, head length, and eye diameter. Which additional keypoints would you add to the label schema to capture these measurements? Consider the trade-off between annotation cost (more keypoints per image means more labeling time) and the scientific value of each additional measurement.

3. Occlusion and visibility scores. Black cod in survey footage frequently swim in loose aggregations, and fins from one individual can overlap the body of another. When a keypoint is occluded, the model may still predict a location for it, but with lower confidence. How would you use the per-keypoint confidence scores returned by .conf to filter out unreliable measurements? At what confidence threshold would you discard a length estimate, and how would you determine that threshold empirically using your validation set?