GLM-4 9B (0414): benchmark results
Provider: Zhipu. Released 2024-06-05. Access: Open.
Unified ELO 1479 ± 45, rank #761 of 1537 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open Portuguese LLM - FaQuAD NLI | 86.86 | Macro F1 (%) | 99.6 |
| Open Portuguese LLM - ASSIN2 RTE | 93.79 | Macro F1 (%) | 92 |
| Open Portuguese LLM - ENEM | 74.32 | Accuracy (%) | 86.7 |
| Open Portuguese LLM - BLUEX | 63.14 | Accuracy (%) | 83.1 |
| Open Portuguese LLM - OAB Exams | 49.66 | Accuracy (%) | 73.9 |
| CommunityBench - Community-Consistent Generation | 176.85 | BTL-Elo rating from majority-voted LLM-judge pairwise compar | 56.2 |
| CommunityBench - Community Identification | 51.25 | Accuracy (%) | 37.5 |
| CommunityBench - Preference Distribution Kendall Tau | 0.06 | Kendall's Tau (-1 to 1) | 31.2 |
| CommunityBench - Preference Identification | 33.12 | Accuracy (%) | 25 |
| CommunityBench - Preference Distribution JSD | 0.17 | Jensen-Shannon Divergence (0-1) | 18.8 |
| ReLE - Education - Primary School Subjects | 30 | Accuracy (%) | 7.1 |
| ReLE - Education - High School Subjects | 27.1 | Accuracy (%) | 6 |
Interactive version: theaggregate.ai/model?slug=glm-4-9b-0414 · How It Works · Data refreshed daily, snapshot 2026-09-25.