GLM-4.5V: benchmark results
Provider: Zhipu. Released 2025-07-28. Access: Open.
Unified ELO 1543 ± 1, rank #469 of 1392 rated models, from 125 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - chinese_homophonic_puns | 1.33 | Dataset z-score | 96.3 |
| FlagEval VQA - Spatial Reasoning | 38.6 | Score | 82.6 |
| MathVision | 65.6 | Overall Accuracy (%) | 78.8 |
| FlagEval VQA - Text Understanding (Chinese) | 75.3 | Score | 78.3 |
| StemBind | 34.6 | F Overall (self-reported) | 69.6 |
| AGC-Bench - grapheval_ai_researcher | 0.44 | Dataset z-score | 67.9 |
| AGC-Bench - analobench | 0.59 | Dataset z-score | 67.3 |
| SciEval - Life Sciences | 54.11 | Life Sciences (%) | 66.7 |
| Math-VR | 49.6 | Overall Answer Correctness (self-reported) | 65.5 |
| Kagi LLM Benchmark | 59.8 | Accuracy (%) | 64.1 |
| Physical AI Bench - Understanding Overall | 59.2 | Overall Score (%) | 62.5 |
| Physical AI Bench Understanding | 59.2 | Overall Score (%) | 62.5 |
Interactive version: theaggregate.ai/model?slug=glm-4-5v · How It Works · Data refreshed daily, snapshot 2026-09-05.