ChartArena - Radar Chart (English): leaderboard

Metric: mAP-high (%) on English radar charts, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 44 models tracked.

Top models

#ModelScore
1GPT-532
2Gemini 3.1 Pro (Preview)31.8
3Kimi K2.5 (Non-reasoning)30.2
4Qwen 3.5 35B A3B (Thinking)25.2
5Qwen 3 VL 235B A22B Instruct23.2
6Qwen 3.5 9B22
7Seed 2.0 Pro (Non-reasoning)21.3
8Gemma 4 31B (IT)20.9
9GLM-4.5V19.7
10MiMo-V2-Omni19.7
11Gemini 2.5 Pro17.5
12Qwen 3 VL 8B Instruct16.8
13Qwen 3 VL 8B (Thinking)14.6
14Qwen 3.5 4B14.1
15InternVL3.5-8B14

Interactive version: theaggregate.ai/benchmark?slug=chartarena-radar-chart-english · How It Works · Data refreshed daily, snapshot 2026-09-29.