Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

3.8 Localizing the OceanCV FirstPass Model to a New Dataset

Open In Colab

Learning Objectives

By the end of this lesson, you will be able to:

  1. Understand how to utilize the OceanCV FirstPass model as a foundational marine ROI detector.

  2. Implement a human-in-the-loop pipeline to adapt a generic “object” detector to specific biological classes.

  3. Fine-tune a YOLO model using domain-specific weights for optimal deep-sea performance.

  4. Analyze model sensitivity across different confidence and IoU thresholds.


Introduction to OceanCV FirstPass

The OceanCV FirstPass model is a broadly trained region-of-interest (ROI) detector built on YOLO and trained on thousands of hours of deep-sea ROV footage. It detects “objects” generically without species-level classification. This lesson shows you how to take that general detector and localize it: adapting its predictions to your specific biological classes through a human-in-the-loop annotation workflow.

This pipeline is fundamental to scalable marine survey work. Rather than starting from scratch with image annotation, you use FirstPass to pre-label candidate objects, then a human expert reviews and re-classifies the detections. The result is a domain-specific model trained from a much smaller hand-labeled dataset.

Step 2: Generate Pre-labels with FirstPass

We load the FirstPass model directly from Hugging Face and run inference on our extracted frames. We use a low confidence threshold (0.10) to capture as many potential objects as possible for human review.


Human-in-the-Loop: Localizing Annotations

Now that the FirstPass model has identified regions of interest, we need to transform these generic “object” detections into biological classes. This is the Human-in-the-Loop phase.

For this example, you will classify the detected objects into four broad ecological tiers:

  1. Sessile Epifauna: Attached organisms (e.g., anemones, sponges).

  2. Motile Epifauna: Bottom-dwelling crawlers (e.g., urchins, sea stars).

  3. Demersal: Swimming animals near the seafloor (e.g., benthic fish).

  4. Planktonic: Organisms drifting in the water column (e.g., jellyfish, larvaceans).

Download dataset_for_labeling.zip, open it in a labeling tool such as Label Studio or CVAT, and assign the correct class to each pre-labeled bounding box. Export the revised annotations in YOLO format. The code below packages the frames and pre-labels for you to download.


Training the Localized Model

After labeling, you are ready to train a model tailored to your specific classes. We will use the OceanCV FirstPass model again, but this time as a pretrained checkpoint to transfer its deep-sea knowledge into our new 4-class classifier.

from ultralytics import YOLO

# Load FirstPass as the starting checkpoint
model = YOLO("OceanCV_FirstPass.pt")

# Train on your localized 4-class dataset
results = model.train(
    data="your_dataset.yaml",
    epochs=100,
    imgsz=1024,
    batch=-1,
    plots=True
)

Threshold Analysis

Once trained, it is vital to understand how the model behaves at different sensitivity levels. Use the code below to visualize predicted results across a grid of Confidence and Intersection over Union (IoU) thresholds.

Ecological Applications: Counting in Zones

Tracking organisms within a defined region (e.g., the bottom third of the frame) allows for standardized ecological surveys. This minimizes noise from drifting plankton in the background and focuses the analysis on the benthic community.

Reflecting on Results

The threshold analysis grid is one of the most diagnostic tools in this pipeline. A low confidence threshold captures more objects but introduces false positives. A high IoU threshold enforces stricter overlap requirements, suppressing duplicate boxes on the same organism.

Reflect on the following:

  • How would you choose a confidence threshold for a real benthic survey where false negatives (missed organisms) are more costly than false positives?

  • What additional ecological tiers would you define for your own survey area?

  • How does pre-labeling with FirstPass change the total annotation time compared to labeling from scratch?