Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

4.3 VLM: MLP and Linear Feature Projection

Open In Colab

A Vision-Language Model connects a visual perception encoder with an autoregressive language model. In foundational architectures such as LLaVA and InternVL, this bridge is a multi-layer perceptron (MLP) or linear projection module. The vision encoder extracts spatial patch embeddings from input imagery. The projection module transforms those visual vectors into the token embedding space of the language model.

This lesson examines the mechanics of MLP projection layers. You will inspect weight dimensions, calculate parameter counts, trace tensor dimension shifts, ingest marine imagery from the LILA BC Community Fish Detection dataset, and execute zero-shot marine organism identification.

Learning Objectives

  • Inspect the 2-layer MLP projection module bridging vision encoder to language model.

  • Trace tensor transformations from visual patch embeddings to language embeddings.

  • Toggle between supported open-weight vision-language models: Qwen2-VL-2B-Instruct and InternVL2-1B.

  • Ingest marine imagery directly from the LILA BC Community Fish Detection dataset with offline fallback support.

  • Execute zero-shot marine organism identification using compact open-weight VLMs.

  • Batch-process marine survey frames and export structured detections to a pandas DataFrame and CSV file.


Part 1 - Setup

1.1 Install Dependencies

This lesson requires transformers, accelerate, pillow, opencv-python, pandas, matplotlib, torchvision, and qwen-vl-utils for vision processing.

1.2 Check GPU and Hardware Fallback

A GPU provides fast model execution. On Google Colab, a free T4 GPU provides 16 GB VRAM. The code detects CUDA automatically and falls back to CPU execution if a GPU is unavailable.

1.3 Hugging Face Authentication Handling

The models used in this lesson (Qwen2-VL-2B-Instruct and InternVL2-1B) are public and ungated. Authentication is optional. If you have a Hugging Face token configured in your environment, the cell below applies it without interrupting execution.


Part 2 - Architecture and Model Loading

2.1 The Two-Layer MLP Projector Concept

A Vision Transformer processes an image into spatial patch embeddings:

Xvis∈RN×DvisX_{vis} \in \mathbb{R}^{N \times D_{vis}}

where NN is the number of patch tokens and DvisD_{vis} is the vision encoder hidden dimension. The language model expects input tokens in its native embedding dimension:

Zllm∈RN×DllmZ_{llm} \in \mathbb{R}^{N \times D_{llm}}

A 2-layer MLP projection module bridges this gap:

H=GELU(XvisW1+b1)H = \text{GELU}(X_{vis} W_1 + b_1)
Zllm=HW2+b2Z_{llm} = H W_2 + b_2

In Qwen2-VL, adjacent 2x2 spatial patch tokens merge prior to projection. Four adjacent tokens of dimension Dvis=1280D_{vis} = 1280 concatenate into a merged vector of dimension Din=5120D_{in} = 5120. The 2-layer MLP maps 5120→5120→15365120 \to 5120 \to 1536, aligning directly with the language model embedding dimension.

In InternVL2-1B, visual patches from an InternViT encoder pass through pixel unshuffle downsampling into an explicit 2-layer MLP module (mlp1). The first linear layer projects the concatenated patch features into intermediate representations with LayerNorm and GELU non-linearities, and the second linear layer projects into the language model embedding dimension (Dllm=896D_{llm} = 896).

2.2 Model Selection and Projector Loading

We provide an interactive model selector toggle supporting both Qwen2-VL-2B-Instruct and InternVL2-1B. You can select either model below to inspect its architecture and run marine organism identification.

2.3 Inspect Projector Weights and Trace Mathematical Transformation

Let us examine the weight shapes of the projection layers and simulate passing a batch of visual tokens through the projector module.


Part 3 - LILA BC Marine Dataset Integration

3.1 Community Fish Detection Dataset Attribution

We integrate imagery from the Community Fish Detection dataset on LILA BC.

The dataset contains underwater camera frames capturing diverse marine teleost species and benthic habitats.

3.2 Download Sample Marine Images

We access individual image blobs directly from Azure storage over HTTP without requiring credentials. We set a default slice of 3 sample images with an adjustable parameter to expand data volume on demand. Network requests are wrapped in exception handlers with synthetic offline fallbacks to ensure uninterrupted execution.

3.3 Visualize Input Marine Frames

We inspect the loaded imagery to observe water clarity, benthic substrate texture, and species visibility.


Part 4 - Inference and Marine Analysis Pipeline

4.1 Define Single-Image Marine Inference Engine

We implement an inference helper that formats the input image and textual query into model inputs. The function supports both Qwen2-VL and InternVL2 architectures.

4.2 Execute Marine Science Analysis Prompt

We run zero-shot inference on the first sample image, requesting species identification, count estimation, and substrate description.


Part 5 - Batch Evaluation and Visualization

5.1 Batch Processing with Latency Tracking and CSV Export

We evaluate all loaded marine images sequentially, recording per-image inference latencies and structured ecological descriptions.

5.2 Visual Summary Display

We display each image alongside its generated identification summary.


Part 6 - Discussion and Marine Science Applications

6.1 Architectural Tradeoffs: MLP vs Resamplers

The linear or 2-layer MLP projection module offers distinct characteristics compared to more complex visual bridges:

  1. Information preservation: Because an MLP projects every visual patch token independently, no visual compression occurs. A 1000-patch image yields 1000 tokens for the language model. This maximizes fine spatial detail, which is critical for identifying small cryptic organisms (such as camouflaged flatfish or juvenile rockfish).

  2. Context length scaling: In high-resolution or multi-frame video applications, preserving all patch tokens rapidly fills the LLM context window. Architecture variants such as Q-Former (Lesson 4.4) and Perceiver Resampler (Lesson 4.6) trade some fine spatial detail for fixed token budgets.

  3. Training efficiency: Projector modules have minimal parameters (often 5 to 15 million parameters), enabling rapid pre-training on aligned image-text pairs while keeping vision and language backbones frozen.

6.2 Human-in-the-Loop Marine Survey Validation

When deploying vision-language models in automated ecological surveys:

  • Use structured JSON prompts to extract standardized Darwin Core biodiversity records.

  • Track confidence metrics and establish human review thresholds for rare or endangered species.

  • Pair zero-shot VLM outputs with traditional supervised object detectors to cross-validate taxonomic classifications.

References