MARINER - VQA: leaderboard

Metric: Accuracy (%) on MARINER maritime visual question answering (mean of the eight task types in perception, space and reasoning; multiple-choice questions generated from verified image metadata and audited by maritime annotators); higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 16 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro73.97
2GPT-4.173.77
3InternVL3-38B69.29
4InternVL3-78B69.05
5GPT-4o68.91
6Qwen 2.5 VL 72B Instruct68.79
7Qwen 2.5 VL 7B Instruct65.63
8InternVL3-8B64.74
9Gemini 2.5 Flash64.52
10Qwen 2.5 VL 32B Instruct63.85
11MiniCPM-V-2.661.57
12InternVL2-8B60.08

Interactive version: theaggregate.ai/benchmark?slug=mariner-vqa · How It Works · Data refreshed daily, snapshot 2026-10-07.