ChartArena - Chinese: leaderboard

Metric: mAP-high (%) averaged over the eight chart types (bar, line, pie, radar, box plot, combination, flowchart, mind map) on Chinese charts, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 42 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)73.2
2Seed 2.0 Pro (Non-reasoning)70.5
3Kimi K2.5 (Non-reasoning)68.1
4Qwen 3.5 35B A3B (Thinking)65.5
5Gemini 2.5 Pro62.4
6Qwen 3 VL 235B A22B Instruct58.4
7Gemma 4 31B (IT)57.6
8MiMo-V2-Omni57
9GLM-4.5V53.9
10Qwen 3 VL 8B Instruct50.4
11Qwen 2.5 VL 72B Instruct50
12Qwen 3 VL 8B (Thinking)49.7
13Qwen 3.5 9B47.7
14Qwen 2.5 VL 32B Instruct47.7
15GPT-545.8

Interactive version: theaggregate.ai/benchmark?slug=chartarena-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.