MARINER - VQA Spatial Relationship: leaderboard

Metric: Accuracy (%) on MARINER maritime visual question answering (spatial relationship questions; multiple-choice questions generated from verified image metadata and audited by maritime annotators); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-4o35.26
2Qwen 2.5 VL 72B Instruct35.26
3Qwen 2.5 VL 32B Instruct35.26
4Gemini 2.5 Flash33.53
5InternVL3-38B33.53
6GPT-4.132.37
7Gemini 2.5 Pro31.79
8InternVL2-8B31.79
9InternVL3-78B30.64
10InternVL3-8B29.48
11Qwen 2.5 VL 7B Instruct27.75
12MiniCPM-V-2.626.01

Interactive version: theaggregate.ai/benchmark?slug=mariner-vqa-spatial-relationship · How It Works · Data refreshed daily, snapshot 2026-10-07.