MARINER - VQA OCR-Integrated QA: leaderboard

Metric: Accuracy (%) on MARINER maritime visual question answering (ship text questions; multiple-choice questions generated from verified image metadata and audited by maritime annotators); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 16 models tracked.

Top models

#ModelScore
1MiniCPM-V-2.691.58
2InternVL3-78B90.1
3InternVL3-38B89.85
4Qwen 2.5 VL 32B Instruct89.11
5Qwen 2.5 VL 72B Instruct88.86
6InternVL2-8B88.61
7Qwen 2.5 VL 7B Instruct88.61
8InternVL3-8B87.38
9Gemini 2.5 Pro86.39
10GPT-4.184.16
11GPT-4o78.22
12Gemini 2.5 Flash58.42

Interactive version: theaggregate.ai/benchmark?slug=mariner-vqa-ocr-integrated-qa · How It Works · Data refreshed daily, snapshot 2026-10-07.