M3-VQA (Gold Entity Names): leaderboard

Metric: IoU accuracy (%) between the predicted and gold answer sets (exact match against Wikidata answer aliases, one-year margin for dates), averaged over the M3-VQA multi-entity, multi-hop knowledge-based visual questions, with the names of the entities behind the gold evidence of every hop added to the input; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-4o51.21
2Qwen 2.5 VL 72B Instruct45.05
3InternVL2.5-78B45.02
4Qwen 2.5 VL 32B Instruct40.39
5MiniCPM-V-2.633.6
6Qwen 2.5 VL 7B Instruct31.83
7Qwen 2 VL 7B Instruct30
8InternVL2.5-2B23.2

Interactive version: theaggregate.ai/benchmark?slug=m3-vqa-gold-entity-names · How It Works · Data refreshed daily, snapshot 2026-10-07.