ChartArena - Combination Chart (Chinese): leaderboard

Metric: mAP-high (%) on Chinese combination charts, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 43 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)70.3
2Kimi K2.5 (Non-reasoning)63.6
3Seed 2.0 Pro (Non-reasoning)62.2
4Gemma 4 31B (IT)59.3
5Qwen 3 VL 235B A22B Instruct58.2
6Gemini 2.5 Pro57.6
7Qwen 3.5 35B A3B (Thinking)56.9
8MiMo-V2-Omni54.7
9GLM-4.5V52.5
10Qwen 2.5 VL 72B Instruct50.5
11Qwen 3.5 9B49.5
12Qwen 3 VL 8B Instruct47.9
13Qwen 2.5 VL 32B Instruct47.3
14Qwen 3 VL 4B Instruct47.1
15Qwen 3.5 4B46.7

Interactive version: theaggregate.ai/benchmark?slug=chartarena-combination-chart-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.