Granite 4.0 Small with guardian: benchmark results
Provider: IBM. Released 2025-10-01. Access: Open.
Unified ELO 1558 ± 1, rank #412 of 1392 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety Anthropic Red Team | 99.9 | LM Evaluated Safety score (%) | 96.5 |
| HELM Safety SimpleSafetyTests | 100 | LM Evaluated Safety score (%) | 86 |
| HELM AIR-Bench | 82.1 | Refusal Rate (%) | 77.9 |
| HELM Safety HarmBench | 86.4 | LM Evaluated Safety score (%) | 69.8 |
| HELM Safety | 91.4 | Mean score (self-reported) | 55.5 |
| HELM Safety BBQ | 90 | BBQ accuracy (%) | 32 |
| HELM Safety XSTest | 80.6 | LM Evaluated Safety score (%) | 3.5 |
Interactive version: theaggregate.ai/model?slug=granite-4-0-small-with-guardian · How It Works · Data refreshed daily, snapshot 2026-09-05.