MARINER - VQA Perspective Estimation: leaderboard

Metric: Accuracy (%) on MARINER maritime visual question answering (vessel perspective questions; multiple-choice questions generated from verified image metadata and audited by maritime annotators); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScore
1GPT-4.172.73
2Gemini 2.5 Pro68.53
3GPT-4o62.7
4Gemini 2.5 Flash59.67
5InternVL3-78B56.41
6Qwen 2.5 VL 72B Instruct56.18
7InternVL3-38B55.94
8Qwen 2.5 VL 7B Instruct51.52
9InternVL3-8B50.58
10Qwen 2.5 VL 32B Instruct49.65
11MiniCPM-V-2.644.29
12InternVL2-8B41.72

Interactive version: theaggregate.ai/benchmark?slug=mariner-vqa-perspective-estimation · How It Works · Data refreshed daily, snapshot 2026-10-07.