OmniMapBench - Multi-Step Reasoning: leaderboard

Metric: Level 3 (multi-step reasoning, 853 questions): strict exact-match accuracy on manually annotated single-choice, multiple-choice and ordering questions over map documents, chain-of-thought prompting (%); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 25 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)62.88
2Qwen 3.5 397B A17B60.37
3Qwen 3.5 Plus57.47
4Kimi K2.556.71
5Qwen 3.5 27B53.97
6Qwen 3.5 122B A10B53.54
7Qwen 3.5 35B A3B49.93
8Qwen 3.5 Flash48.62
9GPT-548.53
10Gemini 2.5 Pro46.31
11GPT-5 Mini45.13
12Gemma 4 31B (IT)40.68
13GPT-4.137.87
14GLM-4.5V35.8
15GPT-4.1 Mini35.64

Interactive version: theaggregate.ai/benchmark?slug=omnimapbench-multi-step-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.