GLM-4.6V (Thinking): benchmark results
Provider: Zhipu. Released 2025-12-08. Access: Open.
Unified ELO 1557 ± 26, rank #836 of 2133 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MultihopSpatial - 1-Hop Exo-Centric | 77.6 | Multiple-choice accuracy (%) on the 1-hop exo-centric questi | 83.3 |
| MultihopSpatial - Acc@50IoU | 34.7 | Acc@50IoU (%): a prediction counts only when the answer is r | 83.3 |
| MultihopSpatial - 2-Hop Exo-Centric | 63.1 | Multiple-choice accuracy (%) on the 2-hop exo-centric questi | 80.6 |
| BenchTable - Tech | 74.3 | Weighted Score (%) | 76.3 |
| MultihopSpatial - 3-Hop Exo-Centric | 46.3 | Multiple-choice accuracy (%) on the 3-hop exo-centric questi | 72.2 |
| MultihopSpatial | 42 | Multiple-choice accuracy (%) over all 4,500 MultihopSpatial | 69.4 |
| BenchTable - STEM | 60.3 | Weighted Score (%) | 63.1 |
| MultihopSpatial - 1-Hop Ego-Centric | 31.2 | Multiple-choice accuracy (%) on the 1-hop ego-centric questi | 62.5 |
| BenchTable | 52.5 | Total Score (%) | 57.3 |
| BenchTable - Reasoning | 40.6 | Weighted Score (%) | 51.2 |
| BenchTable - Utility | 47.4 | Weighted Score (%) | 40.1 |
| 3D Scene Structure - Multi-Swap Rearrangement | 72.33 | Correctness (%, human-scored; 300 scenes; correct only with | 40 |
Interactive version: theaggregate.ai/model?slug=glm-4-6v-thinking · How It Works · Data refreshed daily, snapshot 2026-10-11.