ReasonScape R12 — leaderboard
Reasoning benchmark with 12 cognitive tasks at fixed 16k context. Uses Wilson CIs, truncation penalties, and geometric mean for balanced scoring. Evaluates local/quantized models on structured reasoning.
Metric: ReasonScore. Source: reasonscape.com. Status: saturation imminent. 67 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 4 31B | 927.77 |
| 2 | GLM-4.7 | 894.45 |
| 3 | Qwen 3 235B A22B 2507 Instruct | 841.6 |
| 4 | GLM 4.5 Air | 798.1 |
| 5 | Hunyuan A13B-Instruct | 691.22 |
| 6 | Gemma 3 27B (IT) | 385.1 |
| 7 | Gemma 3 12B (IT) | 286.28 |
| 8 | GLM-4.6V | 145.62 |
Interactive version: theaggregate.ai/benchmark?slug=reasonscape-r12 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.