GLM-4.7 (Reasoning) — benchmark results

GLM-4.7 evaluated with reasoning enabled. Provider: Zhipu. Released 2025-12-22. Access: Open.

Unified ELO 1795 ± 34, rank #144 of 1776 rated models, from 42 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA LiveCodeBench89.42Pass@1 (%)99
AA AIME 202595Accuracy (%)96.7
AA TAU-2 Bench95.91Accuracy (%)96.2
AA MMLU-Pro85.61Accuracy (%)92.7
UGI - Natural Intelligence46.9NatInt Score89.7
UGI Leaderboard49.57UGI Score89.1
UGI - Writing47.38Writing Score87.9
AA GPQA Diamond85.86Accuracy (%)85.8
AA SciCode45.14Accuracy (%)85.7
AA Humanity's Last Exam25.12Accuracy (%)84.6
AA IFBench67.89Accuracy (%)81.4
Artificial Analysis Intelligence Index33.7Intelligence Index81.4

Interactive version: theaggregate.ai/model?slug=glm-4-7-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.