OmniMapBench: leaderboard

Metric: Final score: overall strict exact-match accuracy on manually annotated single-choice, multiple-choice and ordering questions over map documents, chain-of-thought prompting on all 2,096 questions (%); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 25 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)75.03
2Qwen 3.5 397B A17B72.22
3Qwen 3.5 Plus70.8
4Qwen 3.5 27B68.62
5Qwen 3.5 122B A10B68.37
6Kimi K2.567.64
7Qwen 3.5 35B A3B64.75
8Qwen 3.5 Flash64.56
9Gemini 2.5 Pro61.4
10GPT-556.54
11GPT-5 Mini56.11
12Gemma 4 31B (IT)51.67
13GLM-4.5V49.33
14GPT-4.1 Mini45.27
15Claude Sonnet 4.545.04

Interactive version: theaggregate.ai/benchmark?slug=omnimapbench · How It Works · Data refreshed daily, snapshot 2026-09-29.