Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

4.2 LocateAnything FirstPass Localization

Open In Colab

Learning Objectives

By the end of this lesson, you will be able to:

  1. Understand Open-Vocabulary Grounding: Transition from closed-vocabulary detectors (such as the OceanCV FirstPass YOLO model) to open-vocabulary localization using NVIDIA LocateAnything (nvidia/LocateAnything-3B).

  2. Sample ROV Video Transects: Download compressed deep-sea ROV transect footage and extract candidate frames at regular temporal intervals using OpenCV.

  3. Master Non-Expert Morphological Prompting: Formulate visual descriptions in natural language that non-domain experts understand (describing visual morphology such as “large dark purple sea cucumber resting on seafloor” instead of scientific Latin binomials like Paelopatides confundens, “flattened white sea star”, and “translucent swimming jellyfish”).

  4. Parse Structured Spatial Tokens: Extract coordinate tokens (<ref>, <box>) from autoregressive model output into pixel-space coordinates and normalized scales.

  5. Export YOLO Annotations: Convert normalized detections into standard YOLO label format (<class_idx> <x_center> <y_center> <width> <height>).

  6. Package Human-in-the-Loop Datasets: Create a compressed dataset zip archive containing extracted frames and candidate pre-labels for human review in annotation tools such as CVAT or Label Studio.

  7. Evaluate Prompt Variations: Conduct structured exercises comparing prompt phrasing variations across taxonomy, descriptive granularity, substrate context, and distractor disambiguation.


The Paradigm Shift: From Fixed Categories to Open-Vocabulary Grounding

In Lesson 3.8, you utilized the OceanCV FirstPass YOLO detector to find candidate regions of interest. Traditional object detectors rely on a fixed classification layer: they only recognize pre-defined classes learned during supervised training. To detect a new deep-sea taxon, you must collect hundreds of bounding boxes, modify the model head, and retrain network weights.

Vision-Language Models (VLMs) fundamentally change this paradigm. Models such as NVIDIA LocateAnything project visual tokens and natural language embeddings into a shared multimodal space. Instead of being locked to pre-defined classes, LocateAnything locates any target described in free-form natural language text. In this lesson, we replicate the first-pass localization pipeline using LocateAnything, converting raw ROV transect imagery into verified candidate training datasets.

Part 1 - Environment Setup and Hardware Check

We install the necessary libraries for vision-language inference and computer vision processing:

  • transformers >= 4.40.0 and accelerate for model loading and pipeline execution

  • opencv-python-headless for video ingestion and frame sampling

  • pillow and matplotlib for image handling and bounding box rendering

  • pandas for structured tabular inspection of candidate detections

Part 2 - Frame Sampling from ROV Transect Video

Marine surveys record hours of continuous benthic video. Processing every consecutive frame at 30 fps is computationally redundant because adjacent frames show almost identical scenes.

We download a representative ROV benthic transect from Hugging Face (OceanCV/ROVTransectCompressed) and sample candidate frames every NN frames (e.g., sample rate = 32) using OpenCV.

Part 3 - Non-Expert Morphological Prompting vs Scientific Taxonomy

Why Latin Binomials Fail in Web-Scale VLMs

Vision-Language Models learn multimodal associations from hundreds of millions of public image-caption pairs scraped from the internet. These corpora are authored by general internet users who write natural language descriptions of visual scenes (e.g., “purple sea cucumber”, “spiny red crab”, “white sea star”).

Specialized scientific Latin binomials (such as Paelopatides confundens, Benthopecten spinosus, or Bathynomus giganteus) appear with virtually zero frequency in general web datasets. As a result:

  • Asking a VLM for "Paelopatides confundens" produces zero detections or severe hallucinations because the text encoder lacks aligned visual tokens for the taxonomic string.

  • Asking the VLM for "large dark purple sea cucumber resting on seafloor" directly activates the model’s visual-semantic representations of color (dark purple), geometry (elongated sea cucumber), and substrate posture (resting on seafloor).

Guidelines for Formulating Non-Expert Morphological Prompts

When prompting LocateAnything for marine survey localization, think like a visual observer, not a taxonomist:

Scientific Taxonomy (Avoid)Visual Morphology Prompt (Use)Key Visual Cues Encoded
Paelopatides confundens"large dark purple sea cucumber resting on seafloor"Size, deep pigmentation, cylindrical shape, benthic contact
Benthopecten spinosus"flattened white sea star"Flat geometry, stark light color, radiating arm morphology
Scyphozoa / Solmissus"translucent swimming jellyfish"Optical transparency, gelatinous bell shape, pelagic posture
Lithodidae / Paralithodes"red spiny crab on sediment"Prominent coloration, exoskeleton spines, walking legs
Hexactinellida"branching white glass sponge"Porous texture, upright branching silhouette, high contrast

Part 4 - Loading NVIDIA LocateAnything-3B

NVIDIA LocateAnything is a 3-billion-parameter open-vocabulary grounding model. It ingests an image alongside target category strings separated by the </c> token delimiter.

The model autoregressively outputs structured text containing reference tags <ref>...</ref> and bounding coordinate tags <box><x1><y1><x2><y2></box>, where coordinates are normalized to the integer interval [0,1000][0, 1000].

Part 5 - Building Prompt Construction, Parsing, and Visualization Utilities

We implement the core components of the LocateAnything pipeline:

  1. build_messages: Formats chat templates with categories joined by </c>.

  2. parse_detections: Extracts coordinate tokens <box><x1><y1><x2><y2></box> and reference tags <ref> from the model’s textual output.

  3. draw_detections: Renders semi-transparent bounding boxes and labeled tags on candidate frames.

  4. locate_objects: Orchestrates inference with image scaling to prevent out-of-memory errors on GPU instances.

Part 6 - Executing Open-Vocabulary Localization on Transect Imagery

We specify candidate organisms using non-expert morphological natural language descriptions:

  1. "large dark purple sea cucumber resting on seafloor"

  2. "flattened white sea star"

  3. "translucent swimming jellyfish"

  4. "red spiny crab on sediment"

  5. "branching white glass sponge"

We execute localization on our sampled ROV frame, visualize detection boxes, and inspect the detected coordinates in a pandas DataFrame.

Part 7 - Converting Coordinates to Normalized YOLO Format

LocateAnything outputs coordinates in the integer interval [0,1000][0, 1000]. Standard YOLO object detection format requires space-separated text records normalized to [0,1][0, 1]:

⟨class_idx⟩⟨xcenter⟩⟨ycenter⟩⟨width⟩⟨height⟩\langle \text{class\_idx} \rangle \quad \langle x_{\text{center}} \rangle \quad \langle y_{\text{center}} \rangle \quad \langle \text{width} \rangle \quad \langle \text{height} \rangle

The coordinate normalization and bounding box conversion formulas are:

x1′=x11000,y1′=y11000,x2′=x21000,y2′=y21000x_1' = \frac{x_1}{1000}, \quad y_1' = \frac{y_1}{1000}, \quad x_2' = \frac{x_2}{1000}, \quad y_2' = \frac{y_2}{1000}
xcenter=x1′+x2′2,ycenter=y1′+y2′2x_{\text{center}} = \frac{x_1' + x_2'}{2}, \quad y_{\text{center}} = \frac{y_1' + y_2'}{2}
width=x2′−x1′,height=y2′−y1′\text{width} = x_2' - x_1', \quad \text{height} = y_2' - y_1'

We implement detections_to_yolo and write candidate labels to disk.

Part 8 - Packaging Candidate Datasets for Human Review

Zero-shot VLM pre-labels provide an initial pass over raw survey footage. Before training an edge detector (such as YOLOv8 or YOLOv11), a human domain expert must verify the candidate detections, refine bounding box boundaries, and correct false positives.

We package the extracted frames and candidate YOLO label files into a compressed zip archive (dataset_for_labeling.zip). This archive can be uploaded directly to annotation tools such as Label Studio or CVAT.

Part 9 - Structured Exercises: Experimenting with Prompt Phrasing Variations

Because vision-language models ground language through open-vocabulary cross-modal attention, small adjustments in prompt phrasing produce measurable differences in localization recall and precision. Complete the following four structured exercises.


Exercise 1: Scientific Binomial vs Non-Expert Morphological Prompting

Objective: Measure the recall disparity between formal scientific taxonomy and lay visual descriptions.

  1. Select an ROV frame containing a visible sea cucumber or echinoderm.

  2. Run locate_objects using the Latin binomial:

    • Taxonomy prompt: ["Paelopatides confundens"]

  3. Run locate_objects on the same frame using descriptive visual morphology:

    • Morphology prompt: ["large dark purple sea cucumber resting on seafloor"]

  4. Record whether the model successfully detected the organism under each prompt.

  5. Repeat the test for a sea star (["Benthopecten spinosus"] vs ["flattened white sea star"]) and a swimming jellyfish (["Solmissus"] vs ["translucent swimming jellyfish"]).

Discussion Question: Why does the model fail to recognize the Latin name despite being a multi-billion parameter foundation model?


Exercise 2: Descriptive Granularity and Specificity

Objective: Identify the optimal level of morphological detail for candidate pre-labeling. Test three tiers of descriptive granularity on the same frame:

  • Tier 1 (Single generic noun): ["sea star"]

  • Tier 2 (Color + Morphological noun): ["flattened white sea star"]

  • Tier 3 (Compound morphological and postural description): ["flattened white multi-armed sea star crawling on muddy seafloor sediment"]

Evaluation Criteria:

  • Does Tier 1 produce false positives (e.g., detecting pale rocks or sediment patches)?

  • Does Tier 3 overly restrict attention, causing the model to miss the organism?

  • Which tier produces the tightest, most accurate bounding box?


Exercise 3: Substrate and Contextual Conditioning

Objective: Use habitat and substrate context to filter ambiguous detections. In deep-sea video, water column organisms (pelagic) and benthic organisms (attached or crawling on sediment) require different survey treatment. Compare:

  • Unconditioned: ["jellyfish"]

  • Contextually conditioned: ["translucent swimming jellyfish in dark open water column"]

  • Benthic conditioned: ["sessile pale anemone anchored to exposed rocky substrate"]

Observe how explicit contextual keywords guide the model’s spatial attention away from seafloor clutter.


Exercise 4: Disambiguation and Distractor Suppression

Objective: Distinguish between target organisms and visually similar non-biological seafloor features. On deep-sea benthic plains, dark manganese nodules and basalt cobbles frequently mimic holothurians (sea cucumbers):

  • Baseline prompt: ["dark object on seafloor"]

  • Disambiguated organism prompt: ["elongated dark purple sea cucumber with smooth cylindrical body resting on sediment"]

  • Distractor prompt: ["angular dark gray basalt rock"]

Evaluate whether the model correctly assigns separate bounding boxes to the rock and the organism.

Part 10 - Summary and Operational Takeaways

Key Principles of Open-Vocabulary Marine Grounding

  1. Language as an Open Interface: Open-vocabulary models eliminate the constraint of fixed-index classification heads. You can query unannotated transect footage for rare or unmodeled taxa immediately without collecting training sets or fine-tuning weights.

  2. Morphological Prompting is Mandatory: Web-scale models understand colloquial visual descriptions (color, shape, texture, posture) rather than scientific Latin taxonomy. Translating taxonomic taxa into non-expert visual descriptors is a critical domain adaptation skill for marine scientists.

  3. Human-in-the-Loop Pre-Labeling: High-resolution VLMs generate candidate bounding boxes that drastically accelerate dataset creation. Exporting candidates to standardized YOLO format and packaging into zip archives enables human verification in tools like CVAT or Label Studio.

Operational Reflection Questions

  1. When conducting a 100-kilometer ROV survey containing 200,000 frames, why is running an 8GB VLM on every frame impractical, and how does temporal sampling (e.g., every 32nd frame) mitigate this?

  2. How does the quality of human verification in Label Studio affect downstream detector fine-tuning if the VLM produces a 15% false positive rate on manganese nodules?

  3. In what scenarios would you choose to deploy a lightweight YOLOv8 detector over a 3B-parameter vision-language model on an autonomous underwater vehicle (AUV)?