GLM-4.6 (Reasoning) — benchmark results
GLM-4.6 evaluated with reasoning enabled. Provider: Zhipu. Released 2025-09-30. Access: Open.
Unified ELO 1717 ± 16, rank #248 of 1776 rated models, from 58 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI - Writing | 53.31 | Writing Score | 90.4 |
| UGI - Natural Intelligence | 44.94 | NatInt Score | 88.9 |
| AA Omniscience - Software Engineering (SWE) - R | 38 | Accuracy (%) | 84 |
| AA AIME 2025 | 86 | Accuracy (%) | 83.7 |
| AA MMLU-Pro | 82.92 | Accuracy (%) | 82.6 |
| AA Omniscience - Software Engineering (SWE) - Go | 36 | Accuracy (%) | 80.6 |
| AA Omniscience - Software Engineering (SWE) - Rust | 66 | Accuracy (%) | 80.6 |
| AA Omniscience - Science, Engineering & Mathematics | 36.8 | Accuracy (%) | 80.1 |
| AA Global-MMLU-Lite - Korean | 87.67 | Accuracy (%) | 79.1 |
| AA LiveCodeBench | 69.52 | Pass@1 (%) | 78.9 |
| AA Omniscience - Health | 28.2 | Accuracy (%) | 77.2 |
| AA Global-MMLU-Lite - Arabic | 87 | Accuracy (%) | 76.3 |
Interactive version: theaggregate.ai/model?slug=glm-4-6-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.