ReasonScape R12: leaderboard
Reasoning benchmark with 12 cognitive tasks at fixed 16k context. Uses Wilson CIs, truncation penalties, and geometric mean for balanced scoring. Evaluates local/quantized models on structured reasoning.
Metric: ReasonScore. Source: reasonscape.com. Status: saturation imminent. 92 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 4 31B | 927.77 |
| 2 | GLM-4.7 | 894.45 |
| 3 | Qwen 3 235B A22B 2507 Instruct | 841.6 |
| 4 | GLM 4.5 Air | 798.1 |
| 5 | Hunyuan A13B-Instruct | 691.19 |
| 6 | Gemma 4 E4B | 609.7 |
| 7 | GLM-4.7 Flash | 596.84 |
| 8 | GLM-4.6V | 524.07 |
| 9 | Gemma 3 27B (IT) | 385 |
| 10 | Gemma 3 12B (IT) | 284.94 |
Interactive version: theaggregate.ai/benchmark?slug=reasonscape-r12 · How It Works · Data refreshed daily, snapshot 2026-09-05.