DRAGON (Diagrams) - MapWise: leaderboard

Metric: Evidence grounding on maps from MapWise: given the diagram, the question and the verified answer, the model returns bounding boxes for every visual element needed to justify the answer; Grounding IoU (%) between the union of predicted boxes and the union of the human-verified gold evidence boxes, averaged over the EDGE (direct), SAGE (select then ground) and VERGE (verify and refine) prompting strategies; DRAGON test split; temperature 0.2, top-p 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.

Top models

#ModelScore
1Kimi K2.510.8
2Claude Sonnet 4.69.7
3Gemini 3 Pro (Preview)8.5
4Claude Opus 4.68.3
5Llama 4 Maverick Instruct5.2
6Gemma 3 27B (IT)5.1
7Qwen 3.5 35B A3B4.3

Interactive version: theaggregate.ai/benchmark?slug=dragon-diagrams-mapwise · How It Works · Data refreshed daily, snapshot 2026-10-07.