ChartArena - Radar Chart (Chinese): leaderboard

Metric: mAP-high (%) on Chinese radar charts, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 43 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)62.7
2Kimi K2.5 (Non-reasoning)59.7
3Qwen 3.5 35B A3B (Thinking)57.8
4Gemma 4 31B (IT)55.5
5Seed 2.0 Pro (Non-reasoning)54.7
6Gemini 2.5 Pro53
7Qwen 3 VL 235B A22B Instruct52.4
8MiMo-V2-Omni46.1
9Qwen 3.5 9B44.8
10GLM-4.5V43.1
11Qwen 3 VL 8B Instruct42.6
12GPT-541.5
13Qwen 3 VL 8B (Thinking)39.3
14Qwen 2.5 VL 32B Instruct39.2
15Qwen 2.5 VL 72B Instruct38.5

Interactive version: theaggregate.ai/benchmark?slug=chartarena-radar-chart-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.