GLM-4.1V-9B (Thinking) — benchmark results
GLM-4.1V-9B evaluated with thinking enabled. Provider: Zhipu. Released 2025-06-28. Access: Open.
Unified ELO 1428 ± 26, rank #1091 of 1776 rated models, from 129 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval English NLU - SST-5 | 69.38 | Sentiment classification Score (%) | 94.5 |
| EuroEval Italian NLU - Sentipolc16 | 63.7 | Sentiment classification Score (%) | 83.4 |
| EuroEval Faroese NLU - ScaLA FO | 18.25 | Linguistic acceptability Score (%) | 81.5 |
| EuroEval German NLU - Sb10K | 56.64 | Sentiment classification Score (%) | 80.9 |
| EuroEval Finnish NLU - Scandisent FI | 91.54 | Sentiment classification Score (%) | 79.4 |
| EuroEval Spanish NLU - ScaLA ES | 30.64 | Linguistic acceptability Score (%) | 77 |
| EuroEval Danish NLU - Angry Tweets | 53.98 | Sentiment classification Score (%) | 76.1 |
| EuroEval Finnish NLU - ScaLA FI | 31.95 | Linguistic acceptability Score (%) | 76 |
| EuroEval Swedish NLU - Swerec | 77.45 | Sentiment classification Score (%) | 76 |
| EVALITA - text-entailment | 79.79 | CPS | 74.5 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 46.31 | Sentiment classification Score (%) | 73.8 |
| EuroEval English NLU - CoNLL EN | 76.62 | Named entity recognition Score (%) | 72.6 |
Interactive version: theaggregate.ai/model?slug=glm-4-1v-9b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.