GLM-5.1 (Thinking) — benchmark results

Zhipu GLM-5.1 evaluated with thinking enabled. Provider: Zhipu. Released 2026-04-07. Access: Open.

Unified ELO 1812 ± 17, rank #129 of 1776 rated models, from 82 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA TAU-2 Bench97.66Accuracy (%)97.9
Wolfram LLM Benchmarking Project67.8Correct Functionality (%)97
AA IFBench76.26Accuracy (%)95.1
UGI - Writing61.74Writing Score94.5
UGI Leaderboard52.88UGI Score94.5
UGI - Natural Intelligence57.47NatInt Score93.6
Artificial Analysis Intelligence Index40.16Intelligence Index90.4
AA Terminal-Bench Hard43.18Accuracy (%)90.3
Qwen3.7 Launch - IFEval94.5Score (%)90
AA Omniscience1.93Score88.2
AA Humanity's Last Exam28.04Accuracy (%)88.1
AA GPQA Diamond86.77Accuracy (%)87.7

Interactive version: theaggregate.ai/model?slug=glm-5-1-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.