ChartArena - Box Plot (Chinese): leaderboard

Metric: mAP-high (%) on Chinese box plots, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 43 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro (Non-reasoning)55.2
2Qwen 3.5 35B A3B (Thinking)50.6
3Kimi K2.5 (Non-reasoning)47.6
4Gemini 3.1 Pro (Preview)45.2
5GLM-4.5V37.4
6MiMo-V2-Omni30.3
7Qwen 3 VL 4B (Thinking)29
8Qwen 3 VL 8B (Thinking)25.9
9Gemini 2.5 Pro22.1
10Qwen 3 VL 4B Instruct20
11Qwen 3.5 9B18.1
12Qwen 2.5 VL 32B Instruct16.4
13Qwen 2.5 VL 72B Instruct15.3
14Qwen 3 VL 235B A22B Instruct14.1
15GPT-512.8

Interactive version: theaggregate.ai/benchmark?slug=chartarena-box-plot-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.