GLM-4.5: benchmark results

Zhipu's open GLM-4.5 MoE model for reasoning, coding, and agentic tasks (355B total, 32B active). Provider: Zhipu. Released 2025-07-28. Access: Open.

Unified ELO 1611 ± 1, rank #233 of 1392 rated models, from 146 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ZeroEval MATH-50098.2MATH-500 Score93.5
AI for Education Pedagogy - Social studies87.27Accuracy (%)91.4
SEAL - Fortress59.58Score91.1
AI for Education Pedagogy - Technology85.85Accuracy (%)89.5
RAI-Bench - Refusal Rate (General)82Rate (%)85.3
LLM Stats (AIME 2024)91Score (%)84.6
WebCoderBench - Visual Experience87.42Score (%)84.6
LiveMedBench22.46Overall Score (%)83.8
AI for Education SEND81.19Accuracy (%)81
RAI-Bench - RAG Robustness (HY Abstention)88Rate (%)80.4
AI for Education Pedagogy - Primary90.61Accuracy (%)80.3
Enkrypt AI - Bias Risk69.51Risk Score80.3

Interactive version: theaggregate.ai/model?slug=glm-4-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.