GLM-4.1V-9B (Thinking) — benchmark results

GLM-4.1V-9B evaluated with thinking enabled. Provider: Zhipu. Released 2025-06-28. Access: Open.

Unified ELO 1428 ± 26, rank #1091 of 1776 rated models, from 129 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval English NLU - SST-569.38Sentiment classification Score (%)94.5
EuroEval Italian NLU - Sentipolc1663.7Sentiment classification Score (%)83.4
EuroEval Faroese NLU - ScaLA FO18.25Linguistic acceptability Score (%)81.5
EuroEval German NLU - Sb10K56.64Sentiment classification Score (%)80.9
EuroEval Finnish NLU - Scandisent FI91.54Sentiment classification Score (%)79.4
EuroEval Spanish NLU - ScaLA ES30.64Linguistic acceptability Score (%)77
EuroEval Danish NLU - Angry Tweets53.98Sentiment classification Score (%)76.1
EuroEval Finnish NLU - ScaLA FI31.95Linguistic acceptability Score (%)76
EuroEval Swedish NLU - Swerec77.45Sentiment classification Score (%)76
EVALITA - text-entailment79.79CPS74.5
EuroEval Spanish NLU - Sentiment Headlines ES46.31Sentiment classification Score (%)73.8
EuroEval English NLU - CoNLL EN76.62Named entity recognition Score (%)72.6

Interactive version: theaggregate.ai/model?slug=glm-4-1v-9b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.