M3-VQA (Gold Section Evidence): leaderboard

Metric: IoU accuracy (%) between the predicted and gold answer sets (exact match against Wikidata answer aliases, one-year margin for dates), averaged over the M3-VQA multi-entity, multi-hop knowledge-based visual questions, with the Wikipedia sections containing the gold evidence of every hop added to the input; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1InternVL2.5-78B55.16
2GPT-4o53.07
3Qwen 2.5 VL 72B Instruct52.35
4Qwen 2.5 VL 32B Instruct50.02
5Qwen 2.5 VL 7B Instruct42.66
6MiniCPM-V-2.639.11
7Qwen 2 VL 7B Instruct38.92
8InternVL2.5-2B31.94

Interactive version: theaggregate.ai/benchmark?slug=m3-vqa-gold-section-evidence · How It Works · Data refreshed daily, snapshot 2026-10-07.