MapVerse - Reasoning: leaderboard

Metric: Exact match (%) on MapVerse's 304 human-authored multi-step reasoning questions over real-world maps, zero-shot image plus question; higher is better. Source: arxiv.org. Saturation forecast: Around September 2028. 10 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 2.5 VL 7B Instruct32.5#643
2InternVL3-8B25.3#606
3Llama 3.2 11B Instruct24.3#1112
4Mistral Small 3.223.9#587
5aya-vision-8B19.7#1094

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mapverse-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.