Traditional vision-language systems connect an isolated vision encoder to an independent language model through an intermediary adapter. Native multimodal architectures (such as Qwen2-VL) process visual and textual tokens jointly within a unified transformer backbone.
A core innovation in native multimodal processing is 3D Multimodal Rotary Position Embedding (3D M-RoPE). 3D M-RoPE decomposes position IDs into temporal, height, and width components, allowing the model to process arbitrary image resolutions and multi-frame video sequences within a single positional coordinate system.
This lesson explores native multimodal processing. You will inspect special token delimitations, examine 3D positional coordinate calculations, ingest sequential video frames from the Coonamessett River monitoring channel, and execute multi-frame temporal reasoning.
Learning Objectives¶
Understand native multimodal processing versus modular encoder-projector pipelines.
Inspect the 3D M-RoPE positional coordinate mechanism across temporal and spatial axes.
Track special multimodal token markers (
<|vision_start|>,<|vision_end|>).Ingest consecutive video frames from the LILA BC MIT Sea Grant River Herring dataset.
Execute multi-frame temporal reasoning on fish passage and movement dynamics.
Compute inter-frame pixel differences to evaluate motion intensity over time.
Part 1 - Setup¶
1.1 Install Dependencies¶
We install transformers, accelerate, pillow, opencv-python, pandas, matplotlib, and qwen-vl-utils for vision processing.
# Install required dependencies
!pip install -q transformers>=4.40.0 accelerate pillow opencv-python-headless huggingface_hub pandas matplotlib qwen-vl-utils
1.2 Check GPU and Hardware Fallback¶
Multi-image inference requires additional memory for visual token sequences. The script selects CUDA when available and falls back to CPU execution.
import torch
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32
if torch.cuda.is_available():
gpu_name = torch.cuda.get_device_name(0)
vram_gb = torch.cuda.get_device_properties(0).total_memory / 1e9
print(f"[OK] GPU: {gpu_name} | VRAM: {vram_gb:.1f} GB | Precision: {dtype}")
else:
print("[INFO] No GPU detected. Running on CPU with float32 precision.")
1.3 Hugging Face Authentication Handling¶
Qwen2-VL-2B-Instruct is public and ungated. Authentication is optional.
import os
hf_token = os.environ.get("HF_TOKEN", None)
if hf_token:
try:
from huggingface_hub import login
login(token=hf_token)
print("[OK] Hugging Face token applied.")
except Exception as e:
print(f"[INFO] Login skipped: {e}")
else:
print("[INFO] No HF_TOKEN supplied. Using unauthenticated public access.")
Part 2 - Architecture and Model Loading¶
2.1 Native Multimodal Integration and 3D M-RoPE Concept¶
In standard Large Language Models, positional encoding is 1D: a sequence of tokens receives integer positions .
When images and video frames enter the token stream, 1D positions fail to capture 2D spatial adjacency and temporal continuity. 3D M-RoPE resolves this by assigning a 3-tuple coordinate to every token:
Text tokens: receive identical indices across all three axes (), behaving like standard 1D RoPE.
Image tokens: share the same temporal coordinate () but receive distinct spatial grid coordinates:
Video frames: advance the temporal coordinate () while repeating spatial coordinate layouts.
This unified 3D positional formulation allows a single attention mechanism to process text queries, multi-image comparative sets, and continuous video sequences without architectural alterations.
2.2 Model Loading and Special Token Handling¶
We load Qwen2-VL-2B-Instruct and inspect its native multimodal vocabulary tokens.
from transformers import Qwen2VLForConditionalGeneration, AutoProcessor
MODEL_ID = "Qwen/Qwen2-VL-2B-Instruct"
print(f"Loading {MODEL_ID}...")
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = Qwen2VLForConditionalGeneration.from_pretrained(
MODEL_ID,
torch_dtype=dtype,
device_map="auto" if torch.cuda.is_available() else None,
low_cpu_mem_usage=True,
).eval()
if not torch.cuda.is_available():
model = model.to(device)
print(f"[OK] {MODEL_ID} loaded.")
# Inspect special multimodal delimitation tokens
special_tokens = ["<|vision_start|>", "<|vision_end|>", "<|image_pad|>", "<|video_pad|>"]
print("\nMultimodal Special Token IDs:")
for token in special_tokens:
token_id = processor.tokenizer.convert_tokens_to_ids(token)
print(f" {token:<20} -> ID {token_id}")
2.3 Positional Embedding Decomposition¶
We simulate the construction of 3D M-RoPE position grids for a sequence containing text and a 2-frame video clip.
# Simulate positional coordinate generation for a 2-frame video clip
# Suppose each frame is 2x2 patches (4 tokens per frame)
patches_h, patches_w = 2, 2
tokens_per_frame = patches_h * patches_w
num_frames = 2
# Text prompt tokens before and after video
prompt_len = 5
total_tokens = prompt_len + (num_frames * tokens_per_frame)
temporal_pos = []
height_pos = []
width_pos = []
# Initial text tokens: 1D progression across all axes
for i in range(prompt_len):
temporal_pos.append(i)
height_pos.append(i)
width_pos.append(i)
# Video frame tokens
offset = prompt_len
for f_idx in range(num_frames):
for r in range(patches_h):
for c in range(patches_w):
temporal_pos.append(offset + f_idx)
height_pos.append(offset + r)
width_pos.append(offset + c)
print(f"Total simulated token sequence length: {total_tokens}")
print(f"Temporal Position IDs (t): {temporal_pos}")
print(f"Height Position IDs (h): {height_pos}")
print(f"Width Position IDs (w): {width_pos}")
print("\n[OK] 3D M-RoPE coordinate decomposition verified.")
Part 3 - LILA BC Marine Dataset Integration¶
3.1 Coonamessett River Continuous Video Monitoring Attribution¶
We integrate sequential video frames from the MIT Sea Grant River Herring dataset on LILA BC.
Dataset Name: MIT Sea Grant River Herring (Coonamessett River Video Monitoring)
Source Catalog: https://
lila .science /datasets /mit -sea -grant -river -herring/ Attribution: MIT Sea Grant, Town of Falmouth, Massachusetts Division of Marine Fisheries
License: Community Data License Agreement - Permissive (CDLA-Permissive-1.0)
Azure Blob Storage: https://
lilawildlife .blob .core .windows .net /lila -wildlife /mit -river -herring/
The dataset contains continuous underwater camera recordings capturing alewives moving through flumes and observation chambers.
3.2 Download Sequential Video Frames¶
We download consecutive frames representing a short underwater video clip. We set a default slice of 4 consecutive frames with an adjustable parameter to expand data volume on demand.
import urllib.request
import io
from PIL import Image, ImageDraw
BASE_AZURE_URL = "https://lilawildlife.blob.core.windows.net/lila-wildlife/mit-river-herring/Coonamessett/836262/images/default/"
# Consecutive video frames for temporal analysis (10 entries for volume expansion)
SEQUENCE_BLOBS = [
"frame_000000.PNG",
"frame_000001.PNG",
"frame_000002.PNG",
"frame_000003.PNG",
"frame_000004.PNG",
"frame_000005.PNG",
"frame_000006.PNG",
"frame_000007.PNG",
"frame_000008.PNG",
"frame_000009.PNG",
]
# Adjustable data slice parameter (default: 4 consecutive frames)
num_samples = 4 # Increase up to len(SEQUENCE_BLOBS) to download additional frames
sequence_frames = []
for blob_name in SEQUENCE_BLOBS[:num_samples]:
url = BASE_AZURE_URL + blob_name
print(f"Fetching {blob_name}...")
try:
req = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})
with urllib.request.urlopen(req, timeout=15) as resp:
data = resp.read()
pil_img = Image.open(io.BytesIO(data)).convert("RGB")
sequence_frames.append({
"filename": blob_name,
"image": pil_img,
"width": pil_img.width,
"height": pil_img.height,
"bytes": len(data),
"is_fallback": False,
})
print(f" [OK] Loaded {blob_name} ({pil_img.width}x{pil_img.height}, {len(data):,} bytes)")
except Exception as err:
print(f" [WARN] Remote fetch failed for {blob_name} ({err}). Generating synthetic frame fallback.")
fallback_img = Image.new("RGB", (640, 480), color=(15, 60, 85))
draw = ImageDraw.Draw(fallback_img)
draw.rectangle([30, 30, 610, 450], outline=(70, 160, 210), width=3)
draw.text((50, 50), f"Synthetic Sequential Frame: {blob_name}", fill=(215, 240, 255))
sequence_frames.append({
"filename": blob_name,
"image": fallback_img,
"width": 640,
"height": 480,
"bytes": 0,
"is_fallback": True,
})
print(f"\nTotal sequential frames loaded: {len(sequence_frames)}")
3.3 Visualize Video Frame Progression¶
We display the ordered sequence of frames representing the time progression across the river flume.
import matplotlib.pyplot as plt
n_frames = max(1, len(sequence_frames))
fig, axes = plt.subplots(1, n_frames, figsize=(3.5 * n_frames, 3.5))
if n_frames == 1:
axes = [axes]
for idx, (ax, f) in enumerate(zip(axes, sequence_frames)):
ax.imshow(f["image"])
title = f"t={idx} ({f['filename']})" if not f.get("is_fallback") else f"t={idx} (Fallback)"
ax.set_title(title, fontsize=8)
ax.axis("off")
plt.suptitle("Sequential River Monitoring Video Clip", fontsize=11)
plt.tight_layout()
plt.show()
Part 4 - Inference and Temporal Reasoning¶
4.1 Define Multi-Image Inference Engine¶
We pass the complete frame list into a single chat interaction, allowing the model to perform temporal reasoning across consecutive visual states.
from qwen_vl_utils import process_vision_info
@torch.no_grad()
def run_temporal_inference(
images: list[Image.Image],
prompt_text: str,
max_new_tokens: int = 140,
) -> str:
content_list = []
for img in images:
content_list.append({"type": "image", "image": img})
content_list.append({"type": "text", "text": prompt_text})
messages = [
{
"role": "user",
"content": content_list,
}
]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
).to(device)
generated_ids = model.generate(**inputs, max_new_tokens=max_new_tokens)
trimmed_ids = [out[len(inp):] for inp, out in zip(inputs.input_ids, generated_ids)]
response = processor.batch_decode(
trimmed_ids,
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0]
return response.strip()
print("[OK] Temporal multi-image inference engine defined.")
4.2 Multi-Frame Temporal Reasoning¶
We execute temporal inference over the loaded sequence, asking the model to track fish progression across the consecutive frames.
images = [f["image"] for f in sequence_frames]
if not images:
print("No consecutive video frames available (num_samples=0). Increase num_samples to at least 2.")
else:
prompt = (
"These are consecutive video frames recorded at a river fish passage flume. "
"Observe the movement across frames t=0 through t=3. Describe whether any fish enter, "
"advance through the channel, or exit the camera field of view."
)
print(f"Running multi-frame temporal reasoning on {len(images)} consecutive frames...")
response = run_temporal_inference(images, prompt, max_new_tokens=160)
print("\nTemporal Reasoning Output:")
print("-" * 60)
print(response)
print("-" * 60)Part 5 - Batch Evaluation and Temporal Differencing¶
5.1 Inter-Frame Motion Difference Analysis¶
To benchmark the visual changes occurring across the frame sequence, we compute pixel-level absolute difference maps between consecutive frames.
import numpy as np
import pandas as pd
motion_metrics = []
print("Computing pairwise inter-frame difference metrics...")
for i in range(len(sequence_frames) - 1):
f_a = sequence_frames[i]
f_b = sequence_frames[i + 1]
arr_a = np.array(f_a["image"].convert("L"), dtype=np.float32)
arr_b = np.array(f_b["image"].convert("L"), dtype=np.float32)
diff = np.abs(arr_b - arr_a)
mean_motion = float(np.mean(diff))
max_motion = float(np.max(diff))
active_pixel_ratio = float(np.mean(diff > 15.0) * 100.0)
motion_metrics.append({
"transition": f"t{i} -> t{i+1}",
"frame_from": f_a["filename"],
"frame_to": f_b["filename"],
"mean_pixel_diff": round(mean_motion, 2),
"max_pixel_diff": round(max_motion, 2),
"active_pixel_pct": round(active_pixel_ratio, 2),
})
df_motion = pd.DataFrame(motion_metrics)
print(df_motion.to_string(index=False))
5.2 Temporal Feature Comparison Display¶
We visualize the inter-frame difference maps highlighting dynamic regions corresponding to swimming fish and water surface ripples.
if len(sequence_frames) < 2:
print("Insufficient sequential frames to compute inter-frame motion maps (requires at least 2 frames).")
else:
n_trans = len(sequence_frames) - 1
fig, axes = plt.subplots(1, n_trans, figsize=(4 * n_trans, 3.5))
if n_trans == 1:
axes = [axes]
for i, ax in enumerate(axes):
arr_a = np.array(sequence_frames[i]["image"].convert("L"), dtype=np.float32)
arr_b = np.array(sequence_frames[i + 1]["image"].convert("L"), dtype=np.float32)
diff = np.abs(arr_b - arr_a)
im = ax.imshow(diff, cmap="magma")
ax.set_title(f"Motion: t{i}->t{i+1}\nActive: {motion_metrics[i]['active_pixel_pct']}%", fontsize=8)
ax.axis("off")
fig.colorbar(im, ax=ax, fraction=0.046, pad=0.04)
plt.suptitle("Inter-Frame Absolute Difference Maps", fontsize=11)
plt.tight_layout()
plt.show()Part 6 - Discussion and Marine Science Applications¶
6.1 Native Multimodal vs Modular Projector Tradeoffs¶
Joint multimodal processing introduces key operational differences:
Bidirectional cross-attention: In a native multimodal model, later transformer layers allow visual tokens to directly attend to earlier visual tokens from preceding video frames. This enables true temporal tracking within self-attention, rather than relying on separate optical flow or tracking heuristics.
Unified positional space: 3D M-RoPE eliminates the need to flatten video clips into 1D sequences, preserving 2D spatial coherence alongside 1D temporal progression.
Compute demand: Passing multiple consecutive frames through a unified transformer scales sequence length linearly with frame count. For long-term continuous underwater monitoring (e.g., 24-hour flume footage), downsampling strategies or keyframe selection filters are necessary.
6.2 Autonomous Underwater Survey Deployment¶
Deploying native multimodal VLMs on research vessels and coastal stations:
Event-triggered recording: Use lightweight motion differencing (Section 5.1) to detect active passage events, triggering VLM multi-frame reasoning only when organisms cross the observation zone.
Behavioral characterization: Prompt the model to assess schooling cohesion, burst swimming speed, or avoidance behaviors when encountering obstacles or predators.
Edge acceleration: Run quantized models (e.g., 4-bit AWQ or INT8) on edge accelerators (such as NVIDIA Jetson Orin) for continuous coastal monitoring without cloud dependence.
References¶
Qwen2-VL: Wang et al., Qwen2-VL: To See the World More Clearly (2024). https://
arxiv .org /abs /2409 .12191 RoPE: Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding (2024). https://
arxiv .org /abs /2104 .09864 MIT Sea Grant River Herring: https://
lila .science /datasets /mit -sea -grant -river -herring/