DRAGON (Diagrams) - MapIQ: leaderboard

Metric: Evidence grounding on maps from MapIQ: given the diagram, the question and the verified answer, the model returns bounding boxes for every visual element needed to justify the answer; Grounding IoU (%) between the union of predicted boxes and the union of the human-verified gold evidence boxes, averaged over the EDGE (direct), SAGE (select then ground) and VERGE (verify and refine) prompting strategies; DRAGON test split; temperature 0.2, top-p 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.628.4
2Claude Sonnet 4.621.3
3Kimi K2.517.4
4Llama 4 Maverick Instruct6.2
5Qwen 3.5 35B A3B4.7
6Gemma 3 27B (IT)4.5
7Gemini 3 Pro (Preview)4.3

Interactive version: theaggregate.ai/benchmark?slug=dragon-diagrams-mapiq · How It Works · Data refreshed daily, snapshot 2026-10-07.