GLM-5.1: benchmark results
Zhipu GLM-5.1 model for reasoning, coding, and agentic tasks. Provider: Zhipu. Released 2026-03-27. Access: Open.
Unified ELO 1678 ± 1, rank #87 of 1392 rated models, from 417 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - cue_word_story | 1.25 | Dataset z-score | 100 |
| AGC-Bench - simile_generation | 1.9 | Dataset z-score | 100 |
| Are Agents Ready to Teach? A Multi-Stage Bench | 63.8 | Eq. pass (self-reported) | 100 |
| BoxPwnr CTF Bench | 55.47 | Average Platform Completion (%) | 100 |
| GACL - WordMatrix | 68.33 | Normalized Score (0-100) | 100 |
| POLAR-Bench | 94.34 | Overall Mean (self-reported) | 100 |
| SalesBench | 0.51 | 50-Lead Reward | 100 |
| TERMS-Bench | 11.7 | Mean Utility | 100 |
| TimeSage-MT | 40.5 | Overall (self-reported) | 100 |
| MERA - RCB | 63.24 | Accuracy (%) | 99.8 |
| MERA - ruHateSpeech | 93.96 | Accuracy (%) | 99.3 |
| MERA - USE | 77.75 | Grade, normalized (%) | 99 |
Interactive version: theaggregate.ai/model?slug=glm-5-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.