Aerial wildlife censuses and drone-based marine surveys capture high-resolution imagery over vast coastal and pelagic expanses. In traditional vision transformers, images are downscaled to a fixed square (such as 224x224 or 384x384). This downscaling introduces severe geometric aspect-ratio distortion and shrinks small organisms (such as distant seabirds, surfacing cetaceans, or floating waterfowl) into imperceptible sub-pixel artifacts.
Modern vision-language models (such as Qwen2-VL) solve this through dynamic patch resolution and native coordinate grounding. Dynamic resolution partitions images into variable grids of 28x28 patches that preserve the native aspect ratio. Furthermore, the model predicts normalized bounding box coordinates directly in token format:
<|box_start|>(ymin,xmin),(ymax,xmax)<|box_end|>
This lesson explores high-resolution dynamic patching and spatial coordinate grounding. You will inspect dynamic patch grid calculations, build a regular expression coordinate parser, ingest aerial marine wildlife survey imagery from the LILA BC Aerial Seabirds West Africa and Izembek Lagoon Waterfowl datasets, and draw predicted bounding box annotations.
Learning Objectives¶
Understand dynamic patch resolution and aspect-ratio preservation for aerial survey imagery.
Inspect the native Qwen2-VL coordinate grounding token format.
Build a robust parser converting normalized bounding box tokens into pixel coordinates.
Ingest aerial marine survey imagery from LILA BC wildlife datasets.
Execute spatial grounding inference, render detection overlays, and export structured bounding boxes to pandas.
Part 1 - Setup¶
1.1 Install Dependencies¶
We install transformers, accelerate, pillow, opencv-python, pandas, matplotlib, and qwen-vl-utils.
# Install required dependencies
!pip install -q transformers>=4.40.0 accelerate pillow opencv-python-headless huggingface_hub pandas matplotlib qwen-vl-utils
1.2 Check GPU and Hardware Fallback¶
Dynamic resolution processes variable numbers of visual tokens based on image size. A GPU provides optimal throughput, with CPU execution available as a fallback.
import torch
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32
if torch.cuda.is_available():
gpu_name = torch.cuda.get_device_name(0)
vram_gb = torch.cuda.get_device_properties(0).total_memory / 1e9
print(f"[OK] GPU: {gpu_name} | VRAM: {vram_gb:.1f} GB | Precision: {dtype}")
else:
print("[INFO] No GPU detected. Running on CPU with float32 precision.")
1.3 Hugging Face Authentication Handling¶
Qwen2-VL-2B-Instruct is public and ungated. Authentication is optional.
import os
hf_token = os.environ.get("HF_TOKEN", None)
if hf_token:
try:
from huggingface_hub import login
login(token=hf_token)
print("[OK] Hugging Face token applied.")
except Exception as e:
print(f"[INFO] Login skipped: {e}")
else:
print("[INFO] No HF_TOKEN supplied. Using unauthenticated public access.")
Part 2 - Architecture and Model Loading¶
2.1 Dynamic Patch Resolution Mathematics¶
Fixed-resolution models resize an image to a fixed , distorting non-square aerial transects.
Dynamic patch resolution preserves the native aspect ratio by determining the closest valid patch dimensions:
where 28 is the effective patch size (composed of two ViT patches merged spatially).
The resulting total number of visual tokens fed to the LLM is:
The native aspect ratio is maintained within a fraction of a percent, preserving fine organism silhouettes.
2.2 Native Coordinate Grounding Token Format¶
Qwen2-VL expresses spatial bounding boxes natively through special delimiter tokens:
<|box_start|>(ymin,xmin),(ymax,xmax)<|box_end|>
Coordinates are normalized on a integer scale:
ymin: top edge coordinate ()xmin: left edge coordinate ()ymax: bottom edge coordinate ()xmax: right edge coordinate ()
When associated with an object description, the label appears before the bounding box or enclosed in reference tags:
<|object_ref_start|>seabird<|object_ref_end|><|box_start|>(ymin,xmin),(ymax,xmax)<|box_end|>
2.3 Load Model and Processor¶
We load Qwen2-VL-2B-Instruct and its visual processor.
from transformers import Qwen2VLForConditionalGeneration, AutoProcessor
MODEL_ID = "Qwen/Qwen2-VL-2B-Instruct"
print(f"Loading {MODEL_ID}...")
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = Qwen2VLForConditionalGeneration.from_pretrained(
MODEL_ID,
torch_dtype=dtype,
device_map="auto" if torch.cuda.is_available() else None,
low_cpu_mem_usage=True,
).eval()
if not torch.cuda.is_available():
model = model.to(device)
print("[OK] Qwen2-VL-2B-Instruct ready.")
2.4 Test Dynamic Grid Allocation on Aerial Image Dimensions¶
We simulate the dynamic patch calculation for common aerial sensor resolutions (including panoramic drone strips and high-resolution DSLR aerial frames).
test_dimensions = [
(1920, 1080), # Full HD 16:9 aerial video frame
(3840, 2160), # 4K UHD survey transect
(4000, 3000), # 4:3 high-res drone photo
(6000, 2000), # Wide-angle panoramic shoreline survey
(8688, 5792), # USGS Izembek Lagoon DSLR high-res survey frame
]
print("Dynamic Patch Allocation Simulation (Patch Size = 28):")
for w, h in test_dimensions:
# Target pixel budget roughly 256 * 28 * 28 pixels
target_pixels = 256 * 28 * 28
scale = (target_pixels / (w * h)) ** 0.5
h_bar = max(28, int(round(h * scale / 28.0) * 28))
w_bar = max(28, int(round(w * scale / 28.0) * 28))
n_tokens = (h_bar // 28) * (w_bar // 28)
aspect_original = w / h
aspect_resized = w_bar / h_bar
distortion = abs(aspect_resized - aspect_original) / aspect_original * 100
print(f" Input: {w}x{h} (Aspect {aspect_original:.2f}) -> Resized: {w_bar}x{h_bar} (Aspect {aspect_resized:.2f}, Distortion {distortion:.2f}%) -> LLM Tokens: {n_tokens}")
print("\n[OK] Aspect-ratio preservation confirmed.")
Part 3 - LILA BC Marine Dataset Integration¶
3.1 Aerial Seabirds and Waterfowl Datasets Attribution¶
We integrate aerial survey imagery from two LILA BC marine wildlife datasets.
Aerial Seabirds West Africa Dataset:
Source Catalog: https://
lila .science /datasets /aerial -seabirds -west -africa/ Attribution: Kellenberger et al. 2021, Remote Sensing in Ecology and Conservation
License: Community Data License Agreement - Permissive (CDLA-Permissive-1.0)
Azure Blob Storage: https://
lilawildlife .blob .core .windows .net /lila -wildlife /aerial -birds -west -africa/
Izembek Lagoon Waterfowl Imagery:
Source Catalog: https://
lila .science /datasets /izembek -lagoon -waterfowl/ Attribution: U.S. Geological Survey Alaska Science Center
License: Public Domain / CDLA-Permissive
Azure Blob Storage: https://
lilawildlife .blob .core .windows .net /lila -wildlife /izembek -lagoon -birds/
3.2 Download Aerial Marine Survey Samples¶
We download individual sample images directly from Azure storage. We set a default slice of 3 sample frames with an adjustable parameter to expand data volume on demand. Network requests include offline fallback generation.
import urllib.request
import io
from PIL import Image, ImageDraw
# Verified uncompressed sample blobs from LILA BC aerial surveys (10 entries for volume expansion)
AERIAL_BLOBS = [
("aerial_seabirds_west_africa_thumb.jpg", "https://lilawildlife.blob.core.windows.net/lila-wildlife/aerial-birds-west-africa/aerial_seabirds_west_africa_thumb.jpg"),
("2017_Replicate_2017-09-30_Cam1_293A0006.JPG", "https://lilawildlife.blob.core.windows.net/lila-wildlife/izembek-lagoon-birds/images/2017_Replicate_2017-09-30_Cam1_293A0006.JPG"),
("2017_Replicate_2017-09-30_Cam1_293A0007.JPG", "https://lilawildlife.blob.core.windows.net/lila-wildlife/izembek-lagoon-birds/images/2017_Replicate_2017-09-30_Cam1_293A0007.JPG"),
("2017_Replicate_2017-09-30_Cam1_293A0008.JPG", "https://lilawildlife.blob.core.windows.net/lila-wildlife/izembek-lagoon-birds/images/2017_Replicate_2017-09-30_Cam1_293A0008.JPG"),
("2017_Replicate_2017-09-30_Cam1_293A0009.JPG", "https://lilawildlife.blob.core.windows.net/lila-wildlife/izembek-lagoon-birds/images/2017_Replicate_2017-09-30_Cam1_293A0009.JPG"),
("2017_Replicate_2017-09-30_Cam1_293A0010.JPG", "https://lilawildlife.blob.core.windows.net/lila-wildlife/izembek-lagoon-birds/images/2017_Replicate_2017-09-30_Cam1_293A0010.JPG"),
("2017_Replicate_2017-09-30_Cam1_293A0011.JPG", "https://lilawildlife.blob.core.windows.net/lila-wildlife/izembek-lagoon-birds/images/2017_Replicate_2017-09-30_Cam1_293A0011.JPG"),
("2017_Replicate_2017-09-30_Cam1_293A0012.JPG", "https://lilawildlife.blob.core.windows.net/lila-wildlife/izembek-lagoon-birds/images/2017_Replicate_2017-09-30_Cam1_293A0012.JPG"),
("2017_Replicate_2017-09-30_Cam1_293A0013.JPG", "https://lilawildlife.blob.core.windows.net/lila-wildlife/izembek-lagoon-birds/images/2017_Replicate_2017-09-30_Cam1_293A0013.JPG"),
("2017_Replicate_2017-09-30_Cam1_293A0014.JPG", "https://lilawildlife.blob.core.windows.net/lila-wildlife/izembek-lagoon-birds/images/2017_Replicate_2017-09-30_Cam1_293A0014.JPG"),
]
# Adjustable data slice parameter (default: 3)
num_samples = 3 # Increase up to len(AERIAL_BLOBS) to download additional frames
aerial_samples = []
for fname, url in AERIAL_BLOBS[:num_samples]:
print(f"Fetching {fname}...")
try:
req = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})
with urllib.request.urlopen(req, timeout=30) as resp:
data = resp.read()
pil_img = Image.open(io.BytesIO(data)).convert("RGB")
aerial_samples.append({
"filename": fname,
"image": pil_img,
"width": pil_img.width,
"height": pil_img.height,
"bytes": len(data),
"is_fallback": False,
})
print(f" [OK] Loaded {fname} ({pil_img.width}x{pil_img.height}, {len(data):,} bytes)")
except Exception as err:
print(f" [WARN] Remote fetch failed for {fname} ({err}). Generating synthetic fallback.")
fallback_img = Image.new("RGB", (640, 480), color=(50, 70, 90))
draw = ImageDraw.Draw(fallback_img)
draw.rectangle([30, 30, 610, 450], outline=(90, 160, 220), width=3)
draw.text((50, 50), f"Synthetic Aerial Fallback: {fname}", fill=(220, 240, 255))
aerial_samples.append({
"filename": fname,
"image": fallback_img,
"width": 640,
"height": 480,
"bytes": 0,
"is_fallback": True,
})
print(f"\nTotal survey frames loaded: {len(aerial_samples)}")
3.3 Visualize Survey Images¶
We inspect the survey imagery prior to grounding inference.
import matplotlib.pyplot as plt
n_samples = max(1, len(aerial_samples))
fig, axes = plt.subplots(1, n_samples, figsize=(5 * n_samples, 4))
if n_samples == 1:
axes = [axes]
for ax, sample in zip(axes, aerial_samples):
ax.imshow(sample["image"])
title = sample["filename"] if not sample.get("is_fallback") else f"{sample['filename']} (Fallback)"
ax.set_title(title, fontsize=8)
ax.axis("off")
plt.suptitle("Aerial Wildlife Survey Samples", fontsize=11)
plt.tight_layout()
plt.show()
Part 4 - Inference and Detection Pipeline¶
4.1 Define Grounding Coordinate Parser and Inference Engine¶
We implement a regular expression parser that extracts <|box_start|>(ymin,xmin),(ymax,xmax)<|box_end|> tags and converts them to pixel coordinates.
import re
from PIL import ImageDraw, ImageFont
from qwen_vl_utils import process_vision_info
BOX_PATTERN = re.compile(
r"<\|box_start\|>\s*\(\s*([0-9]+(?:\.[0-9]+)?)\s*,\s*([0-9]+(?:\.[0-9]+)?)\s*\)\s*,\s*\(\s*([0-9]+(?:\.[0-9]+)?)\s*,\s*([0-9]+(?:\.[0-9]+)?)\s*\)\s*<\|box_end\|>",
re.IGNORECASE,
)
GROUNDING_PATTERN = re.compile(
r"(?:<\|object_ref_start\|>(.*?)<\|object_ref_end\|>|([a-zA-Z0-9_\s\-]+?))?\s*<\|box_start\|>\s*\(\s*([0-9]+(?:\.[0-9]+)?)\s*,\s*([0-9]+(?:\.[0-9]+)?)\s*\)\s*,\s*\(\s*([0-9]+(?:\.[0-9]+)?)\s*,\s*([0-9]+(?:\.[0-9]+)?)\s*\)\s*<\|box_end\|>"
)
def parse_grounding_output(raw_text: str, img_width: int, img_height: int) -> list[dict]:
'''Extract bounding box coordinates and labels from grounding tokens.'''
detections = []
for m in GROUNDING_PATTERN.finditer(raw_text):
ref_label = m.group(1)
plain_label = m.group(2)
label = ref_label if ref_label is not None else plain_label
if label:
label = label.strip()
label = re.sub(r"^[^a-zA-Z0-9]+", "", label).strip()
if not label:
label = "organism"
ymin_val = float(m.group(3))
xmin_val = float(m.group(4))
ymax_val = float(m.group(5))
xmax_val = float(m.group(6))
max_coord = max(abs(ymin_val), abs(xmin_val), abs(ymax_val), abs(xmax_val))
if 0.0 < max_coord <= 1.05:
scale_x = img_width
scale_y = img_height
else:
scale_x = img_width / 1000.0
scale_y = img_height / 1000.0
x1_val = int(min(xmin_val, xmax_val) * scale_x)
y1_val = int(min(ymin_val, ymax_val) * scale_y)
x2_val = int(max(xmin_val, xmax_val) * scale_x)
y2_val = int(max(ymin_val, ymax_val) * scale_y)
detections.append({
"label": label,
"x1": x1_val,
"y1": y1_val,
"x2": x2_val,
"y2": y2_val,
"norm_coords": [ymin_val, xmin_val, ymax_val, xmax_val],
})
return detections
@torch.no_grad()
def run_grounding_inference(
image: Image.Image,
target_description: str,
max_new_tokens: int = 256,
) -> tuple[str, list[dict]]:
prompt = f"Detect and locate {target_description}. Provide bounding box coordinates in the format <|box_start|>(ymin,xmin),(ymax,xmax)<|box_end|> with labels."
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": prompt},
],
}
]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
).to(device)
generated_ids = model.generate(**inputs, max_new_tokens=max_new_tokens)
trimmed_ids = [out[len(inp):] for inp, out in zip(inputs.input_ids, generated_ids)]
raw_response = processor.batch_decode(
trimmed_ids,
skip_special_tokens=False,
clean_up_tokenization_spaces=False,
)[0]
detections = parse_grounding_output(raw_response, image.width, image.height)
return raw_response, detections
print("[OK] Grounding parser and inference engine defined.")4.2 Execute Grounding Inference on Aerial Sample¶
We execute grounding inference on the aerial survey sample, prompting the model to locate all birds or organisms with bounding boxes.
if not aerial_samples:
print("No aerial survey frames available for grounding (num_samples=0). Increase num_samples to at least 1.")
target = None
raw_text = ""
detections = []
else:
target = aerial_samples[0]
print(f"Running coordinate grounding on {target['filename']}...")
raw_text, detections = run_grounding_inference(target["image"], "all birds, seabirds, or animals in this scene", max_new_tokens=200)
print(f"\nDetections found: {len(detections)}")
print("Raw Model Generation:")
print("-" * 60)
print(raw_text)
print("-" * 60)Part 5 - Batch Evaluation and Visualization¶
5.1 Render Bounding Box Overlays¶
We implement a drawing utility to render predicted bounding boxes and labels onto the survey images.
def draw_bounding_boxes(image: Image.Image, detections: list[dict]) -> Image.Image:
'''Draw colored bounding boxes on a PIL image.'''
canvas = image.copy()
draw = ImageDraw.Draw(canvas)
colors = ["#22c55e", "#06b6d4", "#eab308", "#ec4899", "#8b5cf6"]
for idx, det in enumerate(detections):
color = colors[idx % len(colors)]
raw_x1, raw_y1, raw_x2, raw_y2 = det["x1"], det["y1"], det["x2"], det["y2"]
x1 = min(raw_x1, raw_x2)
x2 = max(raw_x1, raw_x2)
y1 = min(raw_y1, raw_y2)
y2 = max(raw_y1, raw_y2)
draw.rectangle([x1, y1, x2, y2], outline=color, width=3)
label = det["label"]
draw.text((x1 + 4, max(0, y1 - 15)), label, fill=color)
return canvas
if not aerial_samples or target is None:
print("No survey samples available to render bounding boxes (num_samples=0).")
else:
annotated_img = draw_bounding_boxes(target["image"], detections)
fig, axes = plt.subplots(1, 2, figsize=(14, 6))
axes[0].imshow(target["image"])
axes[0].set_title(f"Original: {target['filename']}", fontsize=10)
axes[0].axis("off")
axes[1].imshow(annotated_img)
axes[1].set_title(f"Grounded Detections ({len(detections)} found)", fontsize=10)
axes[1].axis("off")
plt.tight_layout()
plt.show()5.2 Batch Grounding Across All Survey Samples¶
We run grounding inference across all loaded survey samples and compile structured bounding box records into a pandas DataFrame.
import pandas as pd
import time
all_detection_records = []
print("Executing batch grounding across survey samples...")
for idx, sample in enumerate(aerial_samples):
t0 = time.time()
raw, dets = run_grounding_inference(sample["image"], "all organisms, fish, or wildlife", max_new_tokens=150)
elapsed = time.time() - t0
print(f"[{idx+1}/{len(aerial_samples)}] {sample['filename']} processed in {elapsed:.2f}s ({len(dets)} detections)")
for d_idx, d in enumerate(dets):
all_detection_records.append({
"filename": sample["filename"],
"detection_id": d_idx + 1,
"label": d["label"],
"x1_px": d["x1"],
"y1_px": d["y1"],
"x2_px": d["x2"],
"y2_px": d["y2"],
"box_width_px": d["x2"] - d["x1"],
"box_height_px": d["y2"] - d["y1"],
})
df_boxes = pd.DataFrame(all_detection_records)
print("\nStructured Bounding Box Detections:")
if not df_boxes.empty:
print(df_boxes.to_string(index=False))
else:
print("No bounding box detections parsed in sample batch.")
Part 6 - Discussion and Marine Science Applications¶
6.1 Micro-Target Detection and Environmental False Positives¶
Applying vision-language grounding to aerial wildlife surveys introduces unique computer vision challenges:
Micro-target resolution: In wide-angle drone transects (flown at 60 to 120 meters altitude), seabirds or seals may occupy only 15 to 30 pixels. Dynamic patch resolution preserves the native pixel grid, whereas fixed-size downscaling destroys these targets entirely.
Wave froth and glitter false positives: Wind-swept water produces whitecaps, breaking wave foam, and sun glint that closely resemble the white plumage of gulls and terns. Models require sufficient contextual awareness to distinguish wave crest dynamics from biological forms.
Aspect-ratio sensitivity: Pelagic transects often capture wide panoramic strips. Dynamic patching prevents stretching that distorts animal morphology.
6.2 Wildlife Population Census and Ecological Audits¶
Aerial grounding automates animal abundance estimation:
Georeferenced census: Bounding box coordinates map directly to camera GPS coordinates, allowing researchers to convert image counts into spatial density maps (individuals per square kilometer).
Pre-annotation efficiency: Importing VLM-generated bounding box proposals into Label Studio reduces manual annotator workload by 60 to 80 percent.
Human review protocol: Establish confidence thresholds and manual review criteria for ambiguous detections (such as partial submersions or shadowing) before publishing wildlife abundance reports.
References¶
Qwen2-VL: Wang et al., Qwen2-VL: To See the World More Clearly (2024). https://
arxiv .org /abs /2409 .12191 Aerial Seabirds West Africa: Kellenberger et al., Remote Sensing in Ecology and Conservation (2021). https://
lila .science /datasets /aerial -seabirds -west -africa/ Izembek Lagoon Waterfowl: USGS Alaska Science Center. https://
lila .science /datasets /izembek -lagoon -waterfowl/