AIR-BENCH Live: leaderboard
Metric: Mean safety score (x100; %; 2,680 prompts in English, Spanish, Japanese and Portuguese) on regenerated persona-driven attack prompts over 335 level-4 risk categories (one base and one authority-endorsement mutation per category), scored by a GPT-5.4-mini judge as 1 for refusal, 0.5 for partial or evasive, 0 for full compliance. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Haiku 4.5 | 89 |
| 2 | Grok 4.3 | 72 |
| 3 | Llama 3.3 70B | 64 |
| 4 | Gemini 2.5 Pro | 60 |
| 5 | Qwen 3 235B A22B | 60 |
| 6 | Kimi K2 | 58 |
| 7 | GPT-4o | 57 |
| 8 | Gemini 2.5 Flash | 53 |
| 9 | Llama 3 8B | 53 |
| 10 | DeepSeek V3.2 | 52 |
| 11 | DeepSeek R1 | 48 |
| 12 | Mistral Large | 17 |
Interactive version: theaggregate.ai/benchmark?slug=air-bench-live · How It Works · Data refreshed daily, snapshot 2026-09-29.