ReasonScape M12X — leaderboard
12 cognitive domains with 3-degree difficulty scaling. Parametric generators create difficulty manifolds to pinpoint exactly where models break. 100+ models evaluated on 9B+ tokens. Open source.
Metric: ReasonScore (hard). Source: reasonscape.com. Status: saturation imminent. 81 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-4.5 | 529.8 |
| 2 | Qwen 3 Next 80B A3B (Thinking) | 481.4 |
| 3 | Qwen 3 235B A22B FP8 | 473.79 |
| 4 | Hunyuan A13B-Instruct | 467.03 |
| 5 | Qwen 3 VL 32B (Thinking) | 410.44 |
| 6 | Qwen 3 VL 8B (Thinking) | 299.15 |
| 7 | Gemma 3 27B (IT) | 260.24 |
| 8 | Gemma 3 12B (IT) | 196.49 |
| 9 | Qwen 3 VL 4B (Thinking) | 191.77 |
| 10 | OLMo 3 7B (Thinking) | 151.9 |
| 11 | ERNIE-4.5-21B-A3B (Thinking) | 106.43 |
Interactive version: theaggregate.ai/benchmark?slug=reasonscape-m12x · How the rankings work · Data refreshed daily, snapshot 2026-07-22.