ChartArena - Bar Chart (English): leaderboard

Metric: mAP-high (%) on English bar charts, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 44 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)57.9
2Qwen 3.5 35B A3B (Thinking)46.2
3Gemini 2.5 Pro46
4Kimi K2.5 (Non-reasoning)45.2
5Gemma 4 31B (IT)43.7
6Seed 2.0 Pro (Non-reasoning)40.3
7Qwen 3 VL 235B A22B Instruct38.4
8GPT-535.1
9GLM-4.5V33.5
10Qwen 3.5 9B32.5
11Qwen 3 VL 8B (Thinking)32.1
12MiMo-V2-Omni31.1
13Qwen 3 VL 8B Instruct27.5
14Qwen 2.5 VL 32B Instruct27.4
15Qwen 2.5 VL 72B Instruct27.1

Interactive version: theaggregate.ai/benchmark?slug=chartarena-bar-chart-english · How It Works · Data refreshed daily, snapshot 2026-09-29.