GLM-4 9B Chat — benchmark results

Zhipu GLM-4 9B chat model row. Provider: Zhipu. Released 2024-06-05. Access: Open.

Unified ELO 1400 ± 37, rank #1224 of 1776 rated models, from 48 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LIBRA - MatreshkaNames47.33Dataset Total Score (%)92.9
LIBRA - ru2WikiMultihopQA48.78Dataset Total Score (%)92.9
LIBRA - ruBABILongQA322.3Dataset Total Score (%)92.9
LIBRA - ruTREC69.91Dataset Total Score (%)92.9
Fin-Bias98.1Average Herding Score (with rating) (self-reported)88.9
LIBRA - LongContextMultiQ7.75Dataset Total Score (%)78.6
LIBRA - ruSciPassageCount7.5Dataset Total Score (%)78.6
Open PL LLM - RAG69.3Average RAG Score (%)76.9
LogicKor - Multi-Turn7.28Score (0-10)75.4
LogicKor - Reasoning7.07Score (0-10)75.4
Open PL LLM - Generative61.64Average Generative Score (%)74.5
ChineseSafe Benchmark70.96Accuracy (%)74.1

Interactive version: theaggregate.ai/model?slug=glm-4-9b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.