GLM-4.7 (Reasoning): benchmark results
GLM-4.7 evaluated with reasoning enabled. Provider: Zhipu. Released 2025-12-22. Access: Open.
Unified ELO 1624 ± 1, rank #339 of 1761 rated models, from 43 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA LiveCodeBench | 89.42 | Pass@1 (%) | 99.4 |
| AA AIME 2025 | 95 | Accuracy (%) | 97.8 |
| AA TAU-2 Bench | 95.91 | Accuracy (%) | 96.2 |
| AA MMLU-Pro | 85.61 | Accuracy (%) | 92.8 |
| AA Omniscience - Software Engineering (SWE) - HTML | 54 | Accuracy (%) | 91.1 |
| AA Omniscience - Software Engineering (SWE) - PHP | 46 | Accuracy (%) | 90.3 |
| UGI Leaderboard | 49.57 | UGI Score | 89.1 |
| UGI - Natural Intelligence | 46.9 | NatInt Score | 89 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 46.36 | Accuracy (%) | 88.4 |
| AA Omniscience - Software Engineering (SWE) - Dart | 32 | Accuracy (%) | 87.4 |
| UGI - Writing | 47.38 | Writing Score | 87.2 |
| AA Omniscience - Software Engineering (SWE) - C | 55 | Accuracy (%) | 86.7 |
Interactive version: theaggregate.ai/model?slug=glm-4-7-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.