GLM-4.6 (Thinking) — benchmark results

GLM-4.6 evaluated with thinking enabled. Provider: Zhipu. Released 2025-09-30. Access: Open.

Unified ELO 1659 ± 26, rank #348 of 1776 rated models, from 15 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Gorilla API Bench (BFCL)72.38Overall Accuracy (%)96.3
BenchTable72.9Total Score (%)90.4
LLM2014 Code 2025-11 - C#8.38Score84
SpeechMap Compliance78% Requests Completed72.7
LLM2014 Code 2025-1166.81Multi-round Score68
LLM2014 Code 2025-11 - Golang5.7Score64
LLM2014 Code 2025-11 - Java6.87Score64
LLM2014 Code 2025-11 - Python7.22Score56
LLM2014 Logic 2025-1039.97Median Score55.1
LLM2014 Code 2025-11 - TypeScript6.82Score52
LLM2014 Logic 2025-1137.69Median Score51.9
LLM2014 Logic 2025-1233.1Median Score38

Interactive version: theaggregate.ai/model?slug=glm-4-6-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.