GLM-4.6V (Thinking): benchmark results

Provider: Zhipu. Released 2025-12-08. Access: Open.

Unified ELO 1557 ± 26, rank #836 of 2133 rated models, from 19 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MultihopSpatial - 1-Hop Exo-Centric77.6Multiple-choice accuracy (%) on the 1-hop exo-centric questi83.3
MultihopSpatial - Acc@50IoU34.7Acc@50IoU (%): a prediction counts only when the answer is r83.3
MultihopSpatial - 2-Hop Exo-Centric63.1Multiple-choice accuracy (%) on the 2-hop exo-centric questi80.6
BenchTable - Tech74.3Weighted Score (%)76.3
MultihopSpatial - 3-Hop Exo-Centric46.3Multiple-choice accuracy (%) on the 3-hop exo-centric questi72.2
MultihopSpatial42Multiple-choice accuracy (%) over all 4,500 MultihopSpatial 69.4
BenchTable - STEM60.3Weighted Score (%)63.1
MultihopSpatial - 1-Hop Ego-Centric31.2Multiple-choice accuracy (%) on the 1-hop ego-centric questi62.5
BenchTable52.5Total Score (%)57.3
BenchTable - Reasoning40.6Weighted Score (%)51.2
BenchTable - Utility47.4Weighted Score (%)40.1
3D Scene Structure - Multi-Swap Rearrangement72.33Correctness (%, human-scored; 300 scenes; correct only with 40

Interactive version: theaggregate.ai/model?slug=glm-4-6v-thinking · How It Works · Data refreshed daily, snapshot 2026-10-11.