GLM-4.5-Air (Non-reasoning): benchmark results
GLM 4.5 Air evaluated with reasoning disabled. Provider: Zhipu. Released 2025-07-28. Access: Open.
Unified ELO 1529 ± 1, rank #1134 of 3078 rated models, from 150 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Finnish NLU - ScaLA FI | 41.61 | Linguistic acceptability Score (%) | 92.2 |
| EuroEval Spanish NLU - ScaLA ES | 40.78 | Linguistic acceptability Score (%) | 91.2 |
| EuroEval Swedish NLU - Swerec | 78.96 | Sentiment classification Score (%) | 90 |
| EuroEval English Knowledge | 92.31 | Knowledge Average Score (%) | 89.1 |
| EuroEval Italian NLU - Sentipolc16 | 65.51 | Sentiment classification Score (%) | 88.7 |
| EuroEval Norwegian NLU - ScaLA NN | 50.85 | Linguistic acceptability Score (%) | 87.6 |
| EuroEval Portuguese NLU - SST-2 PT | 83.84 | Sentiment classification Score (%) | 86.8 |
| EuroEval German NLU - ScaLA DE | 51.28 | Linguistic acceptability Score (%) | 86.7 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 49.06 | Sentiment classification Score (%) | 86.7 |
| UGI - Writing | 47 | Writing Score | 86.7 |
| EuroEval Norwegian NLU - ScaLA NB | 62.43 | Linguistic acceptability Score (%) | 85.4 |
| EuroEval Norwegian Knowledge - Idioms NO | 33.77 | MCC (x100) | 84.9 |
Interactive version: theaggregate.ai/model?slug=glm-4-5-air-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.