ReasonScape R12: leaderboard

Reasoning benchmark with 12 cognitive tasks at fixed 16k context. Uses Wilson CIs, truncation penalties, and geometric mean for balanced scoring. Evaluates local/quantized models on structured reasoning.

Metric: ReasonScore. Source: reasonscape.com. Status: saturation imminent. 92 models tracked.

Top models

#ModelScore
1Gemma 4 31B927.77
2GLM-4.7894.45
3Qwen 3 235B A22B 2507 Instruct841.6
4GLM 4.5 Air798.1
5Hunyuan A13B-Instruct691.19
6Gemma 4 E4B609.7
7GLM-4.7 Flash596.84
8GLM-4.6V524.07
9Gemma 3 27B (IT)385
10Gemma 3 12B (IT)284.94

Interactive version: theaggregate.ai/benchmark?slug=reasonscape-r12 · How It Works · Data refreshed daily, snapshot 2026-09-05.