MME-Reasoning — leaderboard

1,188-question multimodal reasoning benchmark testing deductive, inductive, and abductive reasoning across calculation, planning, pattern analysis, and spatial-temporal tasks.

Metric: Overall Accuracy (%). Source: alpha-innovator.github.io. Status: years away from saturation. 20 models tracked.

Top models

#ModelScore
1O4 Mini57.2
2Gemini 2.5 Pro53.7
3Claude Sonnet 433
4Claude 3.7 Sonnet32.8
5GPT-4o30.5
6QVQ-72B-Preview28.8
7InternVL3-78B26.5
8Qwen 2 VL 72B24.9
9InternVL3-38B23

Interactive version: theaggregate.ai/benchmark?slug=mme-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.