Learning Objectives¶
By the end of this lesson, you will be able to:
Understand how to utilize the OceanCV FirstPass model as a foundational marine ROI detector.
Implement a human-in-the-loop pipeline to adapt a generic “object” detector to specific biological classes.
Fine-tune a YOLO model using domain-specific weights for optimal deep-sea performance.
Analyze model sensitivity across different confidence and IoU thresholds.
Introduction to OceanCV FirstPass¶
The OceanCV FirstPass model is a broadly trained region-of-interest (ROI) detector built on YOLO and trained on thousands of hours of deep-sea ROV footage. It detects “objects” generically without species-level classification. This lesson shows you how to take that general detector and localize it: adapting its predictions to your specific biological classes through a human-in-the-loop annotation workflow.
This pipeline is fundamental to scalable marine survey work. Rather than starting from scratch with image annotation, you use FirstPass to pre-label candidate objects, then a human expert reviews and re-classifies the detections. The result is a domain-specific model trained from a much smaller hand-labeled dataset.
!pip install ultralytics
!wget -O transect_compressed.mp4 "https://huggingface.co/datasets/OceanCV/ROVTransectCompressed/resolve/main/transect_compressed.mp4?download=true"import os
import cv2
from ultralytics import YOLO
# Create folders for processing
subset_folder = "frames"
os.makedirs(subset_folder, exist_ok=True)
video_path = "transect_compressed.mp4"
cap = cv2.VideoCapture(video_path)
sample_rate = 32 # Extract every 32nd frame
frame_count = 0
print("Extracting frames...")
while cap.isOpened():
ret, frame = cap.read()
if not ret: break
if frame_count % sample_rate == 0:
cv2.imwrite(os.path.join(subset_folder, f"frame_{frame_count}.jpg"), frame)
frame_count += 1
cap.release()
print(f"Extracted {len(os.listdir(subset_folder))} frames.")Step 2: Generate Pre-labels with FirstPass¶
We load the FirstPass model directly from Hugging Face and run inference on our extracted frames. We use a low confidence threshold (0.10) to capture as many potential objects as possible for human review.
# Load the OceanCV FirstPass ROI detector
model = YOLO("https://huggingface.co/OceanCV/OceanCV_FirstPass/resolve/main/OceanCV_FirstPass.pt")
# Run inference to find all potential objects
model.predict(
source=subset_folder,
save_txt=True,
save=True, # Add this to save the annotated images, not strictly necessary but great for stitching a quick video
imgsz=1024,
conf=0.10,
iou=0.5,
project="localization",
name="first_pass"
)Human-in-the-Loop: Localizing Annotations¶
Now that the FirstPass model has identified regions of interest, we need to transform these generic “object” detections into biological classes. This is the Human-in-the-Loop phase.
For this example, you will classify the detected objects into four broad ecological tiers:
Sessile Epifauna: Attached organisms (e.g., anemones, sponges).
Motile Epifauna: Bottom-dwelling crawlers (e.g., urchins, sea stars).
Demersal: Swimming animals near the seafloor (e.g., benthic fish).
Planktonic: Organisms drifting in the water column (e.g., jellyfish, larvaceans).
Download dataset_for_labeling.zip, open it in a labeling tool such as Label Studio or CVAT, and assign the correct class to each pre-labeled bounding box. Export the revised annotations in YOLO format. The code below packages the frames and pre-labels for you to download.
import zipfile
import os
def package_for_labeling(img_folder, label_folder, output_zip):
with zipfile.ZipFile(output_zip, 'w', zipfile.ZIP_DEFLATED) as zipf:
# Add images
for f in os.listdir(img_folder): zipf.write(os.path.join(img_folder, f), arcname=f"images/{f}")
# Add labels
if os.path.exists(label_folder):
for f in os.listdir(label_folder): zipf.write(os.path.join(label_folder, f), arcname=f"labels/{f}")
package_for_labeling("frames", "localization/first_pass/labels", "dataset_for_labeling.zip")
print("Download 'dataset_for_labeling.zip' to import into your labeling environment.")Training the Localized Model¶
After labeling, you are ready to train a model tailored to your specific classes. We will use the OceanCV FirstPass model again, but this time as a pretrained checkpoint to transfer its deep-sea knowledge into our new 4-class classifier.
from ultralytics import YOLO
# Load FirstPass as the starting checkpoint
model = YOLO("OceanCV_FirstPass.pt")
# Train on your localized 4-class dataset
results = model.train(
data="your_dataset.yaml",
epochs=100,
imgsz=1024,
batch=-1,
plots=True
)Threshold Analysis¶
Once trained, it is vital to understand how the model behaves at different sensitivity levels. Use the code below to visualize predicted results across a grid of Confidence and Intersection over Union (IoU) thresholds.
import cv2
import matplotlib.pyplot as plt
from ultralytics import YOLO
model = YOLO("https://huggingface.co/OceanCV/OceanCV_FirstPass/resolve/main/OceanCV_FirstPass.pt")
image_path = "frames/frame_0.jpg"
threshold_settings = [(0.01, 0.5), (0.5, 0.5), (0.1, 0.1), (0.1, 0.9)]
fig, axes = plt.subplots(1, 4, figsize=(24, 6))
for ax, (conf, iou) in zip(axes, threshold_settings):
results = model.predict(image_path, conf=conf, iou=iou, verbose=False)
res_plotted = results[0].plot()
ax.imshow(cv2.cvtColor(res_plotted, cv2.COLOR_BGR2RGB))
ax.set_title(f"Conf: {conf} | IoU: {iou}")
ax.axis('off')
plt.tight_layout()
plt.show()Ecological Applications: Counting in Zones¶
Tracking organisms within a defined region (e.g., the bottom third of the frame) allows for standardized ecological surveys. This minimizes noise from drifting plankton in the background and focuses the analysis on the benthic community.
from ultralytics import solutions
# Example setup for a counting zone
counter = solutions.ObjectCounter(
region=[(0, 700), (1024, 700), (1024, 1024), (0, 1024)], # Bottom region
model="OceanCV_FirstPass.pt",
conf=0.25,
iou=0.45
)
print("Counting zone initialized.")Reflecting on Results¶
The threshold analysis grid is one of the most diagnostic tools in this pipeline. A low confidence threshold captures more objects but introduces false positives. A high IoU threshold enforces stricter overlap requirements, suppressing duplicate boxes on the same organism.
Reflect on the following:
How would you choose a confidence threshold for a real benthic survey where false negatives (missed organisms) are more costly than false positives?
What additional ecological tiers would you define for your own survey area?
How does pre-labeling with FirstPass change the total annotation time compared to labeling from scratch?