DRAGON (Diagrams) - ChartQA: leaderboard

Metric: Evidence grounding on charts from ChartQA: given the diagram, the question and the verified answer, the model returns bounding boxes for every visual element needed to justify the answer; Grounding IoU (%) between the union of predicted boxes and the union of the human-verified gold evidence boxes, averaged over the EDGE (direct), SAGE (select then ground) and VERGE (verify and refine) prompting strategies; DRAGON test split; temperature 0.2, top-p 0.7; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.

Top models

#ModelScore
1Kimi K2.515.1
2Claude Opus 4.611.4
3Gemini 3 Pro (Preview)9.7
4Claude Sonnet 4.67.7
5Qwen 3.5 35B A3B5.7
6Llama 4 Maverick Instruct4.9
7Gemma 3 27B (IT)4.1

Interactive version: theaggregate.ai/benchmark?slug=dragon-diagrams-chartqa · How It Works · Data refreshed daily, snapshot 2026-10-07.