Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

3.18 MarineInst Segmentation

Open In Colab

Overview

MarineInst is a marine-specific instance segmentation model built on top of Segment Anything (SAM), a foundation model released by Meta AI Research in 2023. SAM was trained on a dataset of over one billion masks drawn from 11 million images, giving it a broad understanding of object boundaries across a wide range of visual domains. However, SAM’s training data skews heavily toward terrestrial photography, and the model struggles with phenomena common in marine imagery: organism transparency, disruptive camouflage coloration, and densely cluttered benthic scenes where dozens of species occupy overlapping spatial regions.

MarineInst addresses this by fine-tuning SAM’s components on a curated set of marine imagery, teaching the model to handle the specific edge cases that make underwater segmentation difficult. The core SAM architecture consists of three parts: an image encoder (a large Vision Transformer that converts the input image into a rich spatial feature map), a prompt encoder (which accepts optional user-provided points, boxes, or text to guide segmentation), and a mask decoder (a lightweight transformer that generates one or more candidate masks along with a predicted quality score for each). MarineInst preserves this architecture and refines the weights for the marine domain.

In this lesson you will install MarineInst, download a model checkpoint, run inference on three representative marine images, and inspect the resulting instance masks.

Learning Objectives

By the end of this lesson you will be able to:

  • Explain SAM’s three-component architecture (image encoder, prompt encoder, mask decoder) and describe what each component contributes to the segmentation pipeline.

  • Describe how MarineInst extends SAM through marine-specific fine-tuning and why that matters for transparency and camouflage cases.

  • Run MarineInst inference from the command line with appropriate threshold settings.

  • Interpret instance segmentation overlays and identify where the model succeeds or struggles given scene properties.

Sample Dataset

The three images used in this lesson were selected from the OceanCV image library to test distinct failure modes that plague generic segmentation models in marine contexts:

  • OctopusCrawling.jpg is a textbook camouflage challenge. Octopuses actively match substrate texture and color, so the organism boundary is often a low-contrast seam rather than an obvious edge. A model relying on color contrast will fail here; MarineInst must use shape and contextual cues.

  • SeapigParty.jpg shows a cluster of sea pigs (holothurian scavengers) in contact and partial overlap. The key test is whether the model assigns separate instance masks to touching animals rather than merging them into a single blob. This is the core advantage of instance segmentation over simpler pixel-labeling approaches.

  • WormTower.jpg presents a tubeworm colony: elongated, thin, densely packed structures with fine-grained boundaries. This stresses the mask decoder’s spatial resolution and its ability to trace narrow, curved outlines without merging adjacent tubes.

Running all three images in the same inference pass lets you compare segmentation quality across very different scene types in a single execution.

Downloading the MarineInst Checkpoint

The model weights are stored in a single file: sam_vit_h_4b8939.pth. This is the ViT-H (Vision Transformer Huge) backbone originally trained by Facebook AI Research as part of the Segment Anything project, then adapted by the MarineInst team through continued fine-tuning on marine imagery. The file is approximately 2.4 GB, so a stable connection is required. On Colab, the download typically completes in under two minutes using Google’s network infrastructure.

The filename suffix 4b8939 is a short content hash, not a version number. It uniquely identifies this exact set of weights so you can verify file integrity if the download is interrupted.

Once the checkpoint is on disk, the iou_threshold and sta_threshold parameters at inference time tune how aggressively the model suppresses low-quality mask proposals. These parameters do not change the weights; they filter the decoder’s output candidates before you ever see a mask drawn on screen.

Running MarineInst Inference

The test.py script is MarineInst’s entry point for batch inference. The command-line arguments control the model configuration and output quality filtering:

  • --model_type vit_h selects the ViT-H backbone, which is the largest and most accurate of SAM’s three encoder sizes.

  • --checkpoint points to the downloaded .pth file.

  • --iou_threshold 0.86 instructs the mask decoder to discard any proposed mask whose predicted Intersection over Union score falls below 0.86. Raising this value produces fewer but more confident masks; lowering it recovers more candidates at the cost of more false positives.

  • --sta_threshold 0.92 applies a stability score filter. The stability score measures how consistent the mask boundary is under small perturbations to the input prompt. Values close to 1.0 indicate a crisp, well-defined boundary; values near 0.5 indicate the boundary is ambiguous, which often happens at transparency gradients.

  • --test_img_path points to the folder containing the sample images.

  • --output_dir specifies where the overlay images and raw mask files are written.

These threshold choices were tuned by the MarineInst authors on their marine-specific validation set. You may need to adjust them if you apply the model to very different scene types.

Visualizing Results

The show_results() function reads each original input image and its corresponding MarineInst output from the results directory, then lays them side by side using matplotlib. Each row shows one image: the raw photograph on the left, the mask overlay on the right.

When inspecting the overlays, pay attention to the mask boundaries rather than just the color fills. Good segmentation follows organism outlines even across transparency gradients, where the animal’s interior fades smoothly into the background. If the model has handled camouflage correctly, the mask perimeter will track the animal’s actual edge rather than a high-contrast proxy like a shadow or substrate seam. Look also for whether touching individuals (particularly relevant in the sea pig image) receive separate instance labels, which is the primary advantage of instance segmentation over semantic segmentation.

Reflecting on Results

Take a few minutes to examine the overlay images before moving on.

  1. Which of the three test images produced the cleanest instance masks, and what properties of that image (contrast, background texture, organism shape) explain the difference? Consider how the scene complexity affects both the IoU threshold and the stability score in practice.

  2. MarineInst outputs binary instance masks, one per detected organism. Describe a pipeline that takes those masks, crops the corresponding image regions, and feeds them into a downstream species classifier. What challenges arise when a single mask covers a partially occluded animal?

  3. Beyond the three categories tested here, which additional marine imagery scenarios would most aggressively stress-test MarineInst? Think about gelatinous zooplankton fields, deep-sea bioluminescent scenes, and coral polyp colonies where organism boundaries are functionally meaningless at certain scales.