GLM-130B: benchmark results
Provider: Zhipu. Released 2022-08-04. Access: Open.
Unified ELO 1272 ± 31, rank #2535 of 2656 rated models, from 28 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - CNN/DailyMail | 15.44 | ROUGE-2 (%) | 90.2 |
| HELM Classic - Synthetic Reasoning Natural | 25.35 | F1 (%) | 77.9 |
| HELM Classic - IMDB | 95.47 | Exact Match (%) | 75.8 |
| HELM Classic - NarrativeQA | 70.59 | F1 (%) | 70.8 |
| HELM Classic - NaturalQuestions Open Book | 64.24 | F1 (%) | 70.8 |
| HELM Classic - XSUM | 13.24 | ROUGE-2 (%) | 70.7 |
| HELM Classic - BoolQ | 78.37 | Exact Match (%) | 68.2 |
| HELM Classic - Synthetic Reasoning Abstract | 25.16 | Exact Match (%) | 67.6 |
| HELM Classic - MATH Chain-of-Thought | 5.91 | Equivalent (%) | 57.4 |
| HELM Classic - Dyck | 54.87 | Exact Match (%) | 51.5 |
| HELM Classic - GSM8K | 6.1 | Exact Match (%) | 51.5 |
| HELM Classic - MMLU | 34.4 | Exact Match (%) | 50 |
Interactive version: theaggregate.ai/model?slug=glm-130b · How It Works · Data refreshed daily, snapshot 2026-09-19.