ReasonScape R12 — leaderboard

Reasoning benchmark with 12 cognitive tasks at fixed 16k context. Uses Wilson CIs, truncation penalties, and geometric mean for balanced scoring. Evaluates local/quantized models on structured reasoning.

Metric: ReasonScore. Source: reasonscape.com. Status: saturation imminent. 67 models tracked.

Top models

#ModelScore
1Gemma 4 31B927.77
2GLM-4.7894.45
3Qwen 3 235B A22B 2507 Instruct841.6
4GLM 4.5 Air798.1
5Hunyuan A13B-Instruct691.22
6Gemma 3 27B (IT)385.1
7Gemma 3 12B (IT)286.28
8GLM-4.6V145.62

Interactive version: theaggregate.ai/benchmark?slug=reasonscape-r12 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.