Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

4.7 VLM: High-Resolution Dynamic Patching and Coordinate Grounding

Open In Colab

Aerial wildlife censuses and drone-based marine surveys capture high-resolution imagery over vast coastal and pelagic expanses. In traditional vision transformers, images are downscaled to a fixed square (such as 224x224 or 384x384). This downscaling introduces severe geometric aspect-ratio distortion and shrinks small organisms (such as distant seabirds, surfacing cetaceans, or floating waterfowl) into imperceptible sub-pixel artifacts.

Modern vision-language models (such as Qwen2-VL) solve this through dynamic patch resolution and native coordinate grounding. Dynamic resolution partitions images into variable grids of 28x28 patches that preserve the native aspect ratio. Furthermore, the model predicts normalized bounding box coordinates directly in token format:

<|box_start|>(ymin,xmin),(ymax,xmax)<|box_end|>

This lesson explores high-resolution dynamic patching and spatial coordinate grounding. You will inspect dynamic patch grid calculations, build a regular expression coordinate parser, ingest aerial marine wildlife survey imagery from the LILA BC Aerial Seabirds West Africa and Izembek Lagoon Waterfowl datasets, and draw predicted bounding box annotations.

Learning Objectives

  • Understand dynamic patch resolution and aspect-ratio preservation for aerial survey imagery.

  • Inspect the native Qwen2-VL coordinate grounding token format.

  • Build a robust parser converting normalized bounding box tokens into pixel coordinates.

  • Ingest aerial marine survey imagery from LILA BC wildlife datasets.

  • Execute spatial grounding inference, render detection overlays, and export structured bounding boxes to pandas.


Part 1 - Setup

1.1 Install Dependencies

We install transformers, accelerate, pillow, opencv-python, pandas, matplotlib, and qwen-vl-utils.

1.2 Check GPU and Hardware Fallback

Dynamic resolution processes variable numbers of visual tokens based on image size. A GPU provides optimal throughput, with CPU execution available as a fallback.

1.3 Hugging Face Authentication Handling

Qwen2-VL-2B-Instruct is public and ungated. Authentication is optional.


Part 2 - Architecture and Model Loading

2.1 Dynamic Patch Resolution Mathematics

Fixed-resolution models resize an image (H,W)(H, W) to a fixed (Hfix,Wfix)(H_{fix}, W_{fix}), distorting non-square aerial transects.

Dynamic patch resolution preserves the native aspect ratio by determining the closest valid patch dimensions:

scale=PtargetH×W\text{scale} = \sqrt{\frac{P_{target}}{H \times W}}
Hscaled=round(H×scale28)×28,Wscaled=round(W×scale28)×28H_{scaled} = \text{round}\left(\frac{H \times \text{scale}}{28}\right) \times 28, \quad W_{scaled} = \text{round}\left(\frac{W \times \text{scale}}{28}\right) \times 28

where 28 is the effective patch size (composed of two 14×1414\times 14 ViT patches merged spatially).

The resulting total number of visual tokens fed to the LLM is:

Ntokens=Hscaled28×Wscaled28N_{tokens} = \frac{H_{scaled}}{28} \times \frac{W_{scaled}}{28}

The native aspect ratio is maintained within a fraction of a percent, preserving fine organism silhouettes.

2.2 Native Coordinate Grounding Token Format

Qwen2-VL expresses spatial bounding boxes natively through special delimiter tokens:

<|box_start|>(ymin,xmin),(ymax,xmax)<|box_end|>

Coordinates are normalized on a [0,1000)[0, 1000) integer scale:

  • ymin: top edge coordinate (0≤ymin<10000 \le ymin < 1000)

  • xmin: left edge coordinate (0≤xmin<10000 \le xmin < 1000)

  • ymax: bottom edge coordinate (0<ymax≤10000 < ymax \le 1000)

  • xmax: right edge coordinate (0<xmax≤10000 < xmax \le 1000)

When associated with an object description, the label appears before the bounding box or enclosed in reference tags: <|object_ref_start|>seabird<|object_ref_end|><|box_start|>(ymin,xmin),(ymax,xmax)<|box_end|>

2.3 Load Model and Processor

We load Qwen2-VL-2B-Instruct and its visual processor.

2.4 Test Dynamic Grid Allocation on Aerial Image Dimensions

We simulate the dynamic patch calculation for common aerial sensor resolutions (including panoramic drone strips and high-resolution DSLR aerial frames).


Part 3 - LILA BC Marine Dataset Integration

3.1 Aerial Seabirds and Waterfowl Datasets Attribution

We integrate aerial survey imagery from two LILA BC marine wildlife datasets.

  1. Aerial Seabirds West Africa Dataset:

  1. Izembek Lagoon Waterfowl Imagery:

3.2 Download Aerial Marine Survey Samples

We download individual sample images directly from Azure storage. We set a default slice of 3 sample frames with an adjustable parameter to expand data volume on demand. Network requests include offline fallback generation.

3.3 Visualize Survey Images

We inspect the survey imagery prior to grounding inference.


Part 4 - Inference and Detection Pipeline

4.1 Define Grounding Coordinate Parser and Inference Engine

We implement a regular expression parser that extracts <|box_start|>(ymin,xmin),(ymax,xmax)<|box_end|> tags and converts them to pixel coordinates.

4.2 Execute Grounding Inference on Aerial Sample

We execute grounding inference on the aerial survey sample, prompting the model to locate all birds or organisms with bounding boxes.


Part 5 - Batch Evaluation and Visualization

5.1 Render Bounding Box Overlays

We implement a drawing utility to render predicted bounding boxes and labels onto the survey images.

5.2 Batch Grounding Across All Survey Samples

We run grounding inference across all loaded survey samples and compile structured bounding box records into a pandas DataFrame.


Part 6 - Discussion and Marine Science Applications

6.1 Micro-Target Detection and Environmental False Positives

Applying vision-language grounding to aerial wildlife surveys introduces unique computer vision challenges:

  1. Micro-target resolution: In wide-angle drone transects (flown at 60 to 120 meters altitude), seabirds or seals may occupy only 15 to 30 pixels. Dynamic patch resolution preserves the native pixel grid, whereas fixed-size downscaling destroys these targets entirely.

  2. Wave froth and glitter false positives: Wind-swept water produces whitecaps, breaking wave foam, and sun glint that closely resemble the white plumage of gulls and terns. Models require sufficient contextual awareness to distinguish wave crest dynamics from biological forms.

  3. Aspect-ratio sensitivity: Pelagic transects often capture wide panoramic strips. Dynamic patching prevents stretching that distorts animal morphology.

6.2 Wildlife Population Census and Ecological Audits

Aerial grounding automates animal abundance estimation:

  • Georeferenced census: Bounding box coordinates map directly to camera GPS coordinates, allowing researchers to convert image counts into spatial density maps (individuals per square kilometer).

  • Pre-annotation efficiency: Importing VLM-generated bounding box proposals into Label Studio reduces manual annotator workload by 60 to 80 percent.

  • Human review protocol: Establish confidence thresholds and manual review criteria for ambiguous detections (such as partial submersions or shadowing) before publishing wildlife abundance reports.

References