GLM-4.7 (Reasoning) — benchmark results
GLM-4.7 evaluated with reasoning enabled. Provider: Zhipu. Released 2025-12-22. Access: Open.
Unified ELO 1795 ± 34, rank #144 of 1776 rated models, from 42 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA LiveCodeBench | 89.42 | Pass@1 (%) | 99 |
| AA AIME 2025 | 95 | Accuracy (%) | 96.7 |
| AA TAU-2 Bench | 95.91 | Accuracy (%) | 96.2 |
| AA MMLU-Pro | 85.61 | Accuracy (%) | 92.7 |
| UGI - Natural Intelligence | 46.9 | NatInt Score | 89.7 |
| UGI Leaderboard | 49.57 | UGI Score | 89.1 |
| UGI - Writing | 47.38 | Writing Score | 87.9 |
| AA GPQA Diamond | 85.86 | Accuracy (%) | 85.8 |
| AA SciCode | 45.14 | Accuracy (%) | 85.7 |
| AA Humanity's Last Exam | 25.12 | Accuracy (%) | 84.6 |
| AA IFBench | 67.89 | Accuracy (%) | 81.4 |
| Artificial Analysis Intelligence Index | 33.7 | Intelligence Index | 81.4 |
Interactive version: theaggregate.ai/model?slug=glm-4-7-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.