DRAGON (Diagrams) - AI2D: leaderboard

Metric: Evidence grounding on science diagrams from AI2D: given the diagram, the question and the verified answer, the model returns bounding boxes for every visual element needed to justify the answer; Grounding IoU (%) between the union of predicted boxes and the union of the human-verified gold evidence boxes, averaged over the EDGE (direct), SAGE (select then ground) and VERGE (verify and refine) prompting strategies; DRAGON test split; temperature 0.2, top-p 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.623.8
2Kimi K2.519
3Claude Sonnet 4.615.9
4Gemini 3 Pro (Preview)14.2
5Gemma 3 27B (IT)12.4
6Llama 4 Maverick Instruct9.5
7Qwen 3.5 35B A3B9.2

Interactive version: theaggregate.ai/benchmark?slug=dragon-diagrams-ai2d · How It Works · Data refreshed daily, snapshot 2026-10-07.