DRAGON (Diagrams) - InfographicsVQA: leaderboard
Metric: Evidence grounding on infographics from InfographicsVQA: given the diagram, the question and the verified answer, the model returns bounding boxes for every visual element needed to justify the answer; Grounding IoU (%) between the union of predicted boxes and the union of the human-verified gold evidence boxes, averaged over the EDGE (direct), SAGE (select then ground) and VERGE (verify and refine) prompting strategies; DRAGON test split; temperature 0.2, top-p 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 5.5 |
| 2 | Claude Opus 4.6 | 4.3 |
| 3 | Kimi K2.5 | 3.3 |
| 4 | Claude Sonnet 4.6 | 1.6 |
| 5 | Llama 4 Maverick Instruct | 1.5 |
| 6 | Qwen 3.5 35B A3B | 1.3 |
| 7 | Gemma 3 27B (IT) | 1.1 |
Interactive version: theaggregate.ai/benchmark?slug=dragon-diagrams-infographicsvqa · How It Works · Data refreshed daily, snapshot 2026-10-07.