IBM Granite 4.0 Small — benchmark results
Provider: Other. Released 2025-10-02. Access: Open.
Unified ELO 1565 ± 61, rank #645 of 1839 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety Anthropic Red Team | 99.6 | LM Evaluated Safety score (%) | 83.1 |
| HELM AIR-Bench | 71.6 | Refusal Rate (%) | 60.5 |
| HELM Safety | 91.2 | Mean score (self-reported) | 50.4 |
| HELM Safety HarmBench | 71.1 | LM Evaluated Safety score (%) | 45.3 |
| HELM Safety BBQ | 92.5 | BBQ accuracy (%) | 43 |
| HELM Safety SimpleSafetyTests | 98 | LM Evaluated Safety score (%) | 40.1 |
| HELM Safety XSTest | 94.6 | LM Evaluated Safety score (%) | 36 |
Interactive version: theaggregate.ai/model?slug=ibm-granite-4-0-small · How It Works · Data refreshed daily, snapshot 2026-08-05.