GLM-4.6 (Reasoning): benchmark results
GLM-4.6 evaluated with reasoning enabled. Provider: Zhipu. Released 2025-09-30. Access: Open.
Unified ELO 1591 ± 1, rank #474 of 1761 rated models, from 59 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Omniscience - Software Engineering (SWE) - R | 38 | Accuracy (%) | 93.7 |
| AA Omniscience - Software Engineering (SWE) - Go | 38 | Accuracy (%) | 91.5 |
| AA Omniscience - Software Engineering (SWE) - Swift | 56 | Accuracy (%) | 91.5 |
| UGI - Writing | 53.31 | Writing Score | 89.7 |
| UGI - Natural Intelligence | 44.94 | NatInt Score | 88.2 |
| AA Omniscience - Software Engineering (SWE) - Julia | 28 | Accuracy (%) | 85.2 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 41.82 | Accuracy (%) | 85 |
| AA Omniscience - Software Engineering (SWE) - Java | 25 | Accuracy (%) | 84.6 |
| AA Omniscience - Software Engineering (SWE) - Rust | 62 | Accuracy (%) | 84 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 37.78 | Accuracy (%) | 84 |
| AA AIME 2025 | 86 | Accuracy (%) | 83.7 |
| AA MMLU-Pro | 82.92 | Accuracy (%) | 81.2 |
Interactive version: theaggregate.ai/model?slug=glm-4-6-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.