AIR-BENCH Live (English): leaderboard
Metric: Mean safety score (x100; %; the 670 English prompts) on regenerated persona-driven attack prompts over 335 level-4 risk categories (one base and one authority-endorsement mutation per category), scored by a GPT-5.4-mini judge as 1 for refusal, 0.5 for partial or evasive, 0 for full compliance. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Haiku 4.5 | 91 |
| 2 | Grok 4.3 | 72 |
| 3 | Qwen 3 235B A22B | 67 |
| 4 | Gemini 2.5 Flash | 65 |
| 5 | Llama 3.3 70B | 64 |
| 6 | Gemini 2.5 Pro | 63 |
| 7 | Kimi K2 | 62 |
| 8 | Llama 3 8B | 61 |
| 9 | GPT-4o | 56 |
| 10 | DeepSeek V3.2 | 49 |
| 11 | DeepSeek R1 | 45 |
| 12 | Mistral Large | 20 |
Interactive version: theaggregate.ai/benchmark?slug=air-bench-live-english · How It Works · Data refreshed daily, snapshot 2026-09-29.