MME-Reasoning — leaderboard
1,188-question multimodal reasoning benchmark testing deductive, inductive, and abductive reasoning across calculation, planning, pattern analysis, and spatial-temporal tasks.
Metric: Overall Accuracy (%). Source: alpha-innovator.github.io. Status: years away from saturation. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O4 Mini | 57.2 |
| 2 | Gemini 2.5 Pro | 53.7 |
| 3 | Claude Sonnet 4 | 33 |
| 4 | Claude 3.7 Sonnet | 32.8 |
| 5 | GPT-4o | 30.5 |
| 6 | QVQ-72B-Preview | 28.8 |
| 7 | InternVL3-78B | 26.5 |
| 8 | Qwen 2 VL 72B | 24.9 |
| 9 | InternVL3-38B | 23 |
Interactive version: theaggregate.ai/benchmark?slug=mme-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.