GLM-4.6 — benchmark results

Zhipu's open GLM-4.6 model with a 200K context, improving agentic coding over GLM-4.5. Provider: Zhipu. Released 2025-09-30. Access: Open.

Unified ELO 1647 ± 14, rank #376 of 1776 rated models, from 213 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Diplomacy: Steerability32.5Steerability Score100
AGC-Bench - twistlist1.95Dataset z-score97.6
AGC-Bench - conceptual_design1.41Dataset z-score96.3
AGC-Bench - fig_qa2.38Dataset z-score94.5
AGC-Bench - ocw_connections1.1Dataset z-score93.9
CLEM Hot Air Balloon90.16Game Clemscore (%)91.7
AGC-Bench - metaphoric_analogies1.16Dataset z-score90.2
AGC-Bench - grapheval_iclr0.82Dataset z-score87.8
ConStory-Bench0.53CED errors per 10K words (lower is better)87.5
Diplomacy: Betrayal Tendency85Betrayal Rate (%)87.5
DramaBench93.04Overall Score (%)85.7
AGC-Bench - showerthoughts0.9Dataset z-score85

Interactive version: theaggregate.ai/model?slug=glm-4-6 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.