ChartArena - Chinese: leaderboard
Metric: mAP-high (%) averaged over the eight chart types (bar, line, pie, radar, box plot, combination, flowchart, mind map) on Chinese charts, each model parses the chart image into its native structured format, normalized format-agnostically to (header, entity, value) triples for numeric charts or to labeled graphs and trees for flowcharts and mind maps; mean average precision over similarity thresholds at the high tolerance level, averaged over three visual styles (digital renderings, printed photos, hand-drawn photos); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 42 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 73.2 |
| 2 | Seed 2.0 Pro (Non-reasoning) | 70.5 |
| 3 | Kimi K2.5 (Non-reasoning) | 68.1 |
| 4 | Qwen 3.5 35B A3B (Thinking) | 65.5 |
| 5 | Gemini 2.5 Pro | 62.4 |
| 6 | Qwen 3 VL 235B A22B Instruct | 58.4 |
| 7 | Gemma 4 31B (IT) | 57.6 |
| 8 | MiMo-V2-Omni | 57 |
| 9 | GLM-4.5V | 53.9 |
| 10 | Qwen 3 VL 8B Instruct | 50.4 |
| 11 | Qwen 2.5 VL 72B Instruct | 50 |
| 12 | Qwen 3 VL 8B (Thinking) | 49.7 |
| 13 | Qwen 3.5 9B | 47.7 |
| 14 | Qwen 2.5 VL 32B Instruct | 47.7 |
| 15 | GPT-5 | 45.8 |
Interactive version: theaggregate.ai/benchmark?slug=chartarena-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.