Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Introduction to Vision-Language Models

4.1 Introduction to Vision-Language Models

{: #lesson-vlm-introduction}

Paradigm Shift: Fixed-Vocabulary to Open-Vocabulary Grounding

Marine computer vision workflows historically implement closed-vocabulary object detectors, including YOLO, Faster R-CNN, and RetinaNet. These architectures classify candidate image regions into a fixed index of predetermined category labels. A detector trained on thirty benthic species outputs probability distributions over those thirty specific classes. Extending such models to detect novel taxa, uncataloged developmental stages, or rare morphotypes requires collecting annotated imagery, altering output layer dimensions, and retraining model weights.

Vision-Language Models (VLMs) replace fixed classification heads with multimodal representations that align visual tokens and natural language embeddings within a shared vector space. A vision encoder converts input pixels into spatial feature representations. A multimodal projection layer or cross-attention network maps these spatial representations into the token space of a large language model. The model computes cross-modal similarity between text descriptions and spatial image features. Researchers query unannotated imagery with natural language prompts during inference, achieving open-vocabulary target localization without retraining network parameters.

Architectural Components of Vision-Language Models

Vision-language systems integrate three primary architectural stages:

Visual Encoders: Vision Transformers or convolutional networks divide incoming frames into discrete spatial patches, producing dense sequences of visual feature vectors.

Multimodal Alignment Projectors: Projection layers, including linear projections, multi-layer perceptrons, Querying Transformers, and Perceiver Resamplers, bridge dimensional mismatches between the visual backbone and the language decoder while compressing spatial token sequences.

Language Decoders and Grounding Heads: Autoregressive transformer decoders process combined visual and textual tokens to generate descriptive text or predict spatial bounding box coordinates formatted as numerical tokens.

Grounding models output coordinates directly in normalized token format. Parsers extract these coordinate tokens and convert them to standard pixel-space coordinates and normalized YOLO annotations for downstream workflows.

Practical Grounding in Marine Imagery

Marine survey footage introduces specific optical and linguistic challenges for vision-language models:

Optical Attenuation: Deep-sea and pelagic imagery exhibits wavelength-dependent color absorption, suspended particulate backscatter, and artificial lighting falloff. These optical properties shift visual features relative to terrestrial training datasets.

Taxonomic Vocabulary Mismatch: Web-scale pre-training datasets contain common natural language descriptions, but lack specialized marine taxonomic nomenclature. Vision-language models exhibit low localization recall when prompted with scientific Latin binomials. Prompts specifying observable morphological characteristics, including coloration, geometry, surface texture, and substrate contact, produce reliable spatial grounding.

Deployment Workflows: Autoregressive multimodal models demand significant memory and computational resources. Practical marine monitoring pipelines employ vision-language models as first-pass zero-shot annotators on sampled survey transects, generating candidate bounding boxes for human verification and subsequent training of compact edge detectors.