GLM-4.6V (Non-reasoning) — benchmark results
GLM-4.6V evaluated with reasoning disabled. Provider: Zhipu. Released 2025-12-09. Access: Open.
Unified ELO 1567 ± 19, rank #574 of 1776 rated models, from 56 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI - Natural Intelligence | 32.31 | NatInt Score | 78.8 |
| AA Omniscience - Software Engineering (SWE) - Rust | 56 | Accuracy (%) | 64.3 |
| AA Omniscience - Software Engineering (SWE) - C | 45 | Accuracy (%) | 63.4 |
| AA Omniscience - Software Engineering (SWE) - Swift | 44 | Accuracy (%) | 63.3 |
| AA Omniscience - Software Engineering (SWE) - Julia | 16 | Accuracy (%) | 57.8 |
| UGI - Writing | 35.21 | Writing Score | 56.2 |
| AA Omniscience | -37.87 | Score | 55.1 |
| UGI Leaderboard | 34.72 | UGI Score | 51.3 |
| AA LiveCodeBench | 41.06 | Pass@1 (%) | 49.4 |
| AA MMLU-Pro | 75.16 | Accuracy (%) | 49.4 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 29.09 | Accuracy (%) | 48.9 |
| AA Omniscience - Software Engineering (SWE) | 22.7 | Accuracy (%) | 46.2 |
Interactive version: theaggregate.ai/model?slug=glm-4-6v-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.