ChartArena - Bar Chart (Chinese): leaderboard

Metric: mAP-high (%) on Chinese bar charts, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 43 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)78.7
2Gemini 2.5 Pro76.5
3Seed 2.0 Pro (Non-reasoning)73.3
4Kimi K2.5 (Non-reasoning)70.3
5Gemma 4 31B (IT)68.7
6Qwen 3 VL 235B A22B Instruct67.9
7Qwen 3.5 35B A3B (Thinking)65.3
8GLM-4.5V61.4
9Qwen 2.5 VL 32B Instruct60.1
10Qwen 3 VL 8B Instruct58.6
11MiMo-V2-Omni56.9
12Qwen 3 VL 8B (Thinking)56.5
13Qwen 2.5 VL 72B Instruct53.3
14Qwen 3 VL 4B Instruct52.7
15InternVL3.5-8B52.6

Interactive version: theaggregate.ai/benchmark?slug=chartarena-bar-chart-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.