ChartMuseum — leaderboard

Chart-understanding benchmark reported in Anthropic system cards, with no-tools and Python-tools variants.

Metric: Overall Accuracy (%). Source: chartmuseum-leaderboard.github.io. Status: saturation imminent. 22 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)80.7
2Gemini 3 Flash78.6
3GPT-5.2 (High)77
4Gemma 4 31B (IT)71.1
5Gemma 4 26B A4B (IT)70.6
6GPT-5 Mini (High)63.3
7Gemini 2.5 Pro63
8GPT-5 (High)62.9
9O4 Mini (High)61.5
10O3 (High)60.9
11Claude Opus 4.5 (High)60.7
12Claude Opus 4.159.1
13Claude Sonnet 452.6
14Qwen 3 VL 30B A3B (Thinking)49.7
15Qwen 3 VL 8B (Thinking)44.4

Interactive version: theaggregate.ai/benchmark?slug=chartmuseum · How the rankings work · Data refreshed daily, snapshot 2026-07-22.