ChartArena - Mind Map (English): leaderboard

Metric: mAP-high (%) on English mind maps, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 40 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)86.8
2Seed 2.0 Pro (Non-reasoning)83.1
3Kimi K2.5 (Non-reasoning)80.8
4GPT-576.6
5MiMo-V2-Omni76.6
6Gemma 4 31B (IT)76.2
7Qwen 3.5 35B A3B (Thinking)75.1
8Gemini 2.5 Pro71.7
9Qwen 3 VL 235B A22B Instruct70.8
10Qwen 3 VL 8B Instruct66.4
11GLM-4.5V66.2
12Qwen 3 VL 8B (Thinking)64.8
13Qwen 3.5 9B64.2
14GPT-4o64
15Qwen 2.5 VL 72B Instruct63.8

Interactive version: theaggregate.ai/benchmark?slug=chartarena-mind-map-english · How It Works · Data refreshed daily, snapshot 2026-09-29.