OmniMapBench - Perception: leaderboard

Metric: Level 1 (perception, 441 questions): strict exact-match accuracy on manually annotated single-choice, multiple-choice and ordering questions over map documents, chain-of-thought prompting (%); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 25 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)88.24
2Qwen 3.5 Plus85.32
3Qwen 3.5 397B A17B83.89
4Qwen 3.5 122B A10B83.13
5Qwen 3.5 27B82.93
6Qwen 3.5 35B A3B81.1
7Qwen 3.5 Flash79.27
8Kimi K2.574.77
9Gemini 2.5 Pro73.7
10GLM-4.5V66.67
11GPT-5 Mini65.6
12GPT-563.1
13Intern-S161.22
14Gemma 4 31B (IT)60.77
15GPT-4.1 Mini60.59

Interactive version: theaggregate.ai/benchmark?slug=omnimapbench-perception · How It Works · Data refreshed daily, snapshot 2026-09-29.