M3-VQA (Gold Sentence Evidence): leaderboard

Metric: IoU accuracy (%) between the predicted and gold answer sets (exact match against Wikidata answer aliases, one-year margin for dates), averaged over the M3-VQA multi-entity, multi-hop knowledge-based visual questions, with the gold evidence sentences of every reasoning hop from the linked Wikipedia pages added to the input; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1InternVL2.5-78B58.71
2GPT-4o58.63
3Qwen 2.5 VL 72B Instruct58.41
4Qwen 2.5 VL 32B Instruct53.48
5Qwen 2.5 VL 7B Instruct48.97
6MiniCPM-V-2.646.92
7Qwen 2 VL 7B Instruct41.21
8InternVL2.5-2B34.47

Interactive version: theaggregate.ai/benchmark?slug=m3-vqa-gold-sentence-evidence · How It Works · Data refreshed daily, snapshot 2026-10-07.