MARINER - VQA Attributes: leaderboard

Metric: Accuracy (%) on MARINER maritime visual question answering (vessel attribute questions; multiple-choice questions generated from verified image metadata and audited by maritime annotators); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-4.187.46
2Gemini 2.5 Pro86.02
3GPT-4o82.97
4InternVL3-78B80.84
5Qwen 2.5 VL 72B Instruct80.4
6InternVL3-38B80.36
7Qwen 2.5 VL 32B Instruct75.96
8Gemini 2.5 Flash75.09
9Qwen 2.5 VL 7B Instruct74.26
10InternVL3-8B71.56
11MiniCPM-V-2.668.51
12InternVL2-8B63.81

Interactive version: theaggregate.ai/benchmark?slug=mariner-vqa-attributes · How It Works · Data refreshed daily, snapshot 2026-10-07.