GLM-4.5: benchmark results
Zhipu's open GLM-4.5 MoE model for reasoning, coding, and agentic tasks (355B total, 32B active). Provider: Zhipu. Released 2025-07-28. Access: Open.
Unified ELO 1611 ± 1, rank #233 of 1392 rated models, from 146 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ZeroEval MATH-500 | 98.2 | MATH-500 Score | 93.5 |
| AI for Education Pedagogy - Social studies | 87.27 | Accuracy (%) | 91.4 |
| SEAL - Fortress | 59.58 | Score | 91.1 |
| AI for Education Pedagogy - Technology | 85.85 | Accuracy (%) | 89.5 |
| RAI-Bench - Refusal Rate (General) | 82 | Rate (%) | 85.3 |
| LLM Stats (AIME 2024) | 91 | Score (%) | 84.6 |
| WebCoderBench - Visual Experience | 87.42 | Score (%) | 84.6 |
| LiveMedBench | 22.46 | Overall Score (%) | 83.8 |
| AI for Education SEND | 81.19 | Accuracy (%) | 81 |
| RAI-Bench - RAG Robustness (HY Abstention) | 88 | Rate (%) | 80.4 |
| AI for Education Pedagogy - Primary | 90.61 | Accuracy (%) | 80.3 |
| Enkrypt AI - Bias Risk | 69.51 | Risk Score | 80.3 |
Interactive version: theaggregate.ai/model?slug=glm-4-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.