Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

4.4 VLM: Querying Transformer (Q-Former) Feature Compression

Open In Colab

A Querying Transformer (Q-Former) extracts visual representations of fixed sequence length from variable-resolution image patch grids. In foundational vision-language architectures such as BLIP-2 and InstructBLIP, the Q-Former bridges an invariant visual encoder with a frozen large language model.

This lesson implements a functional PyTorch Q-Former module from scratch. You will inspect cross-attention mechanics, compress variable patch counts into a fixed budget of 32 query tokens, ingest river herring migration imagery from the LILA BC MIT Sea Grant dataset, and run migration count analysis.

Learning Objectives

  • Implement a functional PyTorch Querying Transformer (Q-Former) module with learned query embeddings and cross-attention blocks.

  • Compress variable visual token sequences into a fixed budget of 32 query tokens.

  • Trace cross-attention feature extraction from spatial patch representations.

  • Ingest river herring migration frames from the LILA BC MIT Sea Grant dataset with offline fallback handling.

  • Execute zero-shot fish migration analysis with compact open-weight VLMs.

  • Export structured observation records to a pandas DataFrame and CSV file.


Part 1 - Setup

1.1 Install Dependencies

We install transformers, accelerate, pillow, opencv-python, pandas, matplotlib, and qwen-vl-utils for vision processing.

1.2 Check GPU and Hardware Fallback

A GPU provides accelerated tensor operations. The module automatically selects CUDA when available and falls back to CPU execution.

1.3 Hugging Face Authentication Handling

Public model weights load without authentication. You may supply a personal HF token if desired.


Part 2 - Architecture and Module Implementation

2.1 The Q-Former Compression Mechanism

Standard ViT encoders produce N=(H/14)×(W/14)N = (H / 14) \times (W / 14) patch tokens. For high-resolution imagery, NN easily exceeds 1000 tokens, straining language model context limits.

The Q-Former employs a set of KK learned query embeddings:

Q0∈RK×DqQ_0 \in \mathbb{R}^{K \times D_q}

Through LL transformer blocks, the queries interact with each other via self-attention and extract visual features from encoder tokens Xvis∈RN×DvisX_{vis} \in \mathbb{R}^{N \times D_{vis}} via cross-attention:

S=SelfAttention(LN(Q))+QS = \text{SelfAttention}(\text{LN}(Q)) + Q
C=CrossAttention(Query=LN(S),Key=Xvis,Value=Xvis)+SC = \text{CrossAttention}(\text{Query}=\text{LN}(S), \text{Key}=X_{vis}, \text{Value}=X_{vis}) + S
Qnext=FFN(LN(C))+CQ_{next} = \text{FFN}(\text{LN}(C)) + C

Regardless of whether N=256N = 256, 576, or 2048, the output always comprises exactly KK query tokens, providing fixed-length visual conditioning.

2.2 Functional PyTorch Q-Former Implementation

We implement the complete Q-Former module using native PyTorch multihead attention and linear layers.

2.3 Verify Multi-Head Attention and Fixed 32-Token Compression

We test the module with different visual input sizes (representing low, medium, and high resolution frames) to confirm that the output always compresses into exactly 32 tokens.

2.4 Load Downstream Vision-Language Model

We load Qwen2-VL-2B-Instruct to execute downstream zero-shot reasoning on the loaded survey imagery.


Part 3 - LILA BC Marine Dataset Integration

3.1 MIT Sea Grant River Herring Dataset Attribution

We integrate camera frames from the MIT Sea Grant River Herring dataset on LILA BC.

The dataset monitors anadromous river herring (alewife and blueback herring) migrating upstream through fish ladders and restoration channels in coastal Massachusetts.

3.2 Download Sample River Herring Migration Frames

We access sequential video frames directly from Azure storage. We set a default slice of 3 frames with an adjustable parameter to expand data volume on demand. Network requests include offline fallback generation.

3.3 Visualize Video Frames

We inspect the river herring migration frames to assess flume illumination, turbidity, and fish silhouettes.


Part 4 - Inference and Ecological Analysis Pipeline

4.1 Define River Herring Inference Engine

We define an inference routine specialized for river flume monitoring, instructing the model to observe fish silhouettes, count passing individuals, and note swimming orientation.

4.2 Passage Count and Orientation Analysis

We analyze the first video frame, requesting detection of river herring, swimming direction (upstream vs downstream), and water clarity assessment.


Part 5 - Batch Evaluation and Attention Pooling

5.1 Batch Evaluation Across Sequential Frames

We process consecutive frames to evaluate temporal migration patterns and export results to CSV.

5.2 Simulated Spatial Attention Pooling Heatmap Display

To understand how the Q-Former query tokens aggregate spatial information, we simulate an attention pooling heatmap overlaid onto the migration imagery.


Part 6 - Discussion and Marine Science Applications

6.1 Information Bottleneck Tradeoffs

The Q-Former introduces an intentional information bottleneck:

  • Compression ratio: Compressing 1024 patch tokens into 32 query tokens achieves a 32x token reduction. This drastically reduces the computational burden on the language model attention matrix, which scales quadratically with sequence length.

  • Selective attention: Through cross-attention, learnable query embeddings dynamically allocate attention to salient targets (such as passing fish silhouettes) while ignoring static channel walls and murky water backgrounds.

  • Fine-grained loss: For small organisms that span only a few pixels, aggressive token compression can cause spatial detail loss. High-resolution grounding (Lesson 4.7) explores alternative architectures that retain dynamic spatial resolution.

6.2 Autonomous Video Monitoring in River Restoration

Automating passage counts along migratory corridors:

  • Real-time run estimates: Continuous video indexing allows fisheries biologists to calculate seasonal run totals without manually watching hundreds of hours of video.

  • Species discrimination: Vision-language models differentiate river herring from non-target species (such as white suckers, sea lamprey, or striped bass) using subtle morphological cues.

  • Turbidity robustness: Multi-head cross-attention learns to pool features across low-contrast water frames where simple background subtraction fails.

References