GLM-4.5-Air-FP8 — benchmark results
Provider: Zhipu. Released 2025-07-28. Access: Open.
Unified ELO 1673 ± 63, rank #376 of 1839 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety XSTest | 98.6 | LM Evaluated Safety score (%) | 94.8 |
| HELM Safety BBQ | 97.8 | BBQ accuracy (%) | 93 |
| HELM Safety SimpleSafetyTests | 99 | LM Evaluated Safety score (%) | 57.6 |
| Vectara Hallucination Leaderboard | 90.7 | Factual Consistency Rate (%) | 55.3 |
| HELM Safety Anthropic Red Team | 99 | LM Evaluated Safety score (%) | 53.5 |
| HELM Safety | 90.1 | Mean score (self-reported) | 40.7 |
| ForecastBench | 62.4 | Overall Score (higher is better) | 38.3 |
| HELM AIR-Bench | 57.1 | Refusal Rate (%) | 30.2 |
| HELM Safety HarmBench | 56.1 | LM Evaluated Safety score (%) | 19.8 |
Interactive version: theaggregate.ai/model?slug=glm-4-5-air-fp8 · How It Works · Data refreshed daily, snapshot 2026-08-05.