GLM-4.6 — benchmark results
Zhipu's open GLM-4.6 model with a 200K context, improving agentic coding over GLM-4.5. Provider: Zhipu. Released 2025-09-30. Access: Open.
Unified ELO 1647 ± 14, rank #376 of 1776 rated models, from 213 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Diplomacy: Steerability | 32.5 | Steerability Score | 100 |
| AGC-Bench - twistlist | 1.95 | Dataset z-score | 97.6 |
| AGC-Bench - conceptual_design | 1.41 | Dataset z-score | 96.3 |
| AGC-Bench - fig_qa | 2.38 | Dataset z-score | 94.5 |
| AGC-Bench - ocw_connections | 1.1 | Dataset z-score | 93.9 |
| CLEM Hot Air Balloon | 90.16 | Game Clemscore (%) | 91.7 |
| AGC-Bench - metaphoric_analogies | 1.16 | Dataset z-score | 90.2 |
| AGC-Bench - grapheval_iclr | 0.82 | Dataset z-score | 87.8 |
| ConStory-Bench | 0.53 | CED errors per 10K words (lower is better) | 87.5 |
| Diplomacy: Betrayal Tendency | 85 | Betrayal Rate (%) | 87.5 |
| DramaBench | 93.04 | Overall Score (%) | 85.7 |
| AGC-Bench - showerthoughts | 0.9 | Dataset z-score | 85 |
Interactive version: theaggregate.ai/model?slug=glm-4-6 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.