ChartArena - Combination Chart (English): leaderboard

Metric: mAP-high (%) on English combination charts, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 44 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)39.7
2Kimi K2.5 (Non-reasoning)33.6
3Seed 2.0 Pro (Non-reasoning)32.4
4Qwen 3.5 35B A3B (Thinking)31.5
5Qwen 3 VL 235B A22B Instruct29.1
6Gemini 2.5 Pro28.7
7Gemma 4 31B (IT)22.7
8GLM-4.5V21.2
9MiMo-V2-Omni19.4
10Qwen 3.5 9B16.8
11Qwen 3 VL 8B (Thinking)16
12Qwen 2.5 VL 72B Instruct14.3
13GPT-514.2
14Qwen 3 VL 8B Instruct13.2
15InternVL3.5-8B11.3

Interactive version: theaggregate.ai/benchmark?slug=chartarena-combination-chart-english · How It Works · Data refreshed daily, snapshot 2026-09-29.