GLM-4 9B Chat — benchmark results
Zhipu GLM-4 9B chat model row. Provider: Zhipu. Released 2024-06-05. Access: Open.
Unified ELO 1400 ± 37, rank #1224 of 1776 rated models, from 48 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LIBRA - MatreshkaNames | 47.33 | Dataset Total Score (%) | 92.9 |
| LIBRA - ru2WikiMultihopQA | 48.78 | Dataset Total Score (%) | 92.9 |
| LIBRA - ruBABILongQA3 | 22.3 | Dataset Total Score (%) | 92.9 |
| LIBRA - ruTREC | 69.91 | Dataset Total Score (%) | 92.9 |
| Fin-Bias | 98.1 | Average Herding Score (with rating) (self-reported) | 88.9 |
| LIBRA - LongContextMultiQ | 7.75 | Dataset Total Score (%) | 78.6 |
| LIBRA - ruSciPassageCount | 7.5 | Dataset Total Score (%) | 78.6 |
| Open PL LLM - RAG | 69.3 | Average RAG Score (%) | 76.9 |
| LogicKor - Multi-Turn | 7.28 | Score (0-10) | 75.4 |
| LogicKor - Reasoning | 7.07 | Score (0-10) | 75.4 |
| Open PL LLM - Generative | 61.64 | Average Generative Score (%) | 74.5 |
| ChineseSafe Benchmark | 70.96 | Accuracy (%) | 74.1 |
Interactive version: theaggregate.ai/model?slug=glm-4-9b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.