ChartArena - Mind Map (Chinese): leaderboard

Metric: mAP-high (%) on Chinese mind maps, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 38 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro (Non-reasoning)85.8
2Gemini 3.1 Pro (Preview)85.2
3Kimi K2.5 (Non-reasoning)79.4
4Qwen 3.5 35B A3B (Thinking)70.9
5Gemini 2.5 Pro67.1
6Qwen 3 VL 235B A22B Instruct65.2
7MiMo-V2-Omni64.6
8Gemma 4 31B (IT)61
9Qwen 2.5 VL 72B Instruct55
10Qwen 3 VL 8B Instruct54.6
11Qwen 3.5 9B54.5
12GLM-4.5V43.7
13Qwen 3 VL 8B (Thinking)42.3
14Qwen 2.5 VL 32B Instruct40.6
15Qwen 3.5 4B39.6

Interactive version: theaggregate.ai/benchmark?slug=chartarena-mind-map-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.