GLM-4.6V-Flash: benchmark results
Provider: Zhipu. Access: Open.
Unified ELO 1703 ± 26, rank #321 of 2653 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GEB-Bench - Story-Theorem Matching | 47.1 | Accuracy (%; 4-way, chance 25%) | 68.1 |
| GEB-Bench - Scene-Theorem Matching | 52.9 | Accuracy (%; 4-way, chance 25%) | 61.5 |
| OCRBench-V2 (zh) | 59.5 | Score (self-reported) | 60 |
| GEB-Bench - Theorem Identification | 76 | Accuracy (%; 25-way, chance 4%) | 59.7 |
| GEB-Bench - Scene-Story Matching | 32.9 | Accuracy (%; 4-way, chance 25%) | 59.6 |
| OCRBench v2 | 59.5 | Average (self-reported) | 57.7 |
| TRAP | 19.7 | Overall HM (self-reported) | 57.1 |
| GEB-Bench - Story Identification | 38.8 | Accuracy (%; 25-way, chance 4%) | 55.6 |
| GEB-Bench - Adversarial Matching | 72.4 | Accuracy (%; 4-way, chance 25%) | 53.8 |
| GEB-Bench - Average (Vision Models) | 53 | Macro-average accuracy over all eight tasks (%) | 53.8 |
| OCRBench-V2 (en) | 59 | Score (self-reported) | 52 |
| SciTrue - Claim Validation | 79.6 | F1 (%) | 50 |
Interactive version: theaggregate.ai/model?slug=glm-4-6v-flash · How It Works · Data refreshed daily, snapshot 2026-09-21.