DRAGON (Diagrams) - MapIQ: leaderboard
Metric: Evidence grounding on maps from MapIQ: given the diagram, the question and the verified answer, the model returns bounding boxes for every visual element needed to justify the answer; Grounding IoU (%) between the union of predicted boxes and the union of the human-verified gold evidence boxes, averaged over the EDGE (direct), SAGE (select then ground) and VERGE (verify and refine) prompting strategies; DRAGON test split; temperature 0.2, top-p 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 28.4 |
| 2 | Claude Sonnet 4.6 | 21.3 |
| 3 | Kimi K2.5 | 17.4 |
| 4 | Llama 4 Maverick Instruct | 6.2 |
| 5 | Qwen 3.5 35B A3B | 4.7 |
| 6 | Gemma 3 27B (IT) | 4.5 |
| 7 | Gemini 3 Pro (Preview) | 4.3 |
Interactive version: theaggregate.ai/benchmark?slug=dragon-diagrams-mapiq · How It Works · Data refreshed daily, snapshot 2026-10-07.