GLM 4.5 Air (Thinking): benchmark results
Provider: Zhipu. Released 2025-07-28. Access: Open.
Unified ELO 1544 ± 1, rank #1021 of 3078 rated models, from 146 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Lithuanian NLU - WikiANN LT | 70.41 | Named entity recognition Score (%) | 96.1 |
| EuroEval Latvian NLU - Fullstack NER LV | 71.41 | Named entity recognition Score (%) | 95.7 |
| EuroEval English Knowledge | 94.32 | Knowledge Average Score (%) | 95 |
| EuroEval Icelandic NLU - MIM-GOLD NER | 75.93 | Named entity recognition Score (%) | 94.4 |
| EuroEval Italian NLU - MultiNERD IT | 82.52 | Named entity recognition Score (%) | 93.2 |
| EuroEval Dutch NLU - CoNLL NL | 72.82 | Named entity recognition Score (%) | 92.4 |
| EuroEval Norwegian NLU - NorNE NN | 79.48 | Named entity recognition Score (%) | 91.9 |
| EuroEval Portuguese NLU - HAREM | 58.3 | Named entity recognition Score (%) | 91.9 |
| EuroEval German NLU - Sb10K | 58.92 | Sentiment classification Score (%) | 91.7 |
| EuroEval French Knowledge | 77.58 | Knowledge Average Score (%) | 91.5 |
| EuroEval Spanish NLU - CoNLL ES | 76.98 | Named entity recognition Score (%) | 91.5 |
| EuroEval Icelandic Knowledge | 44.5 | Knowledge Average Score (%) | 91.4 |
Interactive version: theaggregate.ai/model?slug=glm-4-5-air-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.