AIR-BENCH Live: leaderboard

Metric: Mean safety score (x100; %; 2,680 prompts in English, Spanish, Japanese and Portuguese) on regenerated persona-driven attack prompts over 335 level-4 risk categories (one base and one authority-endorsement mutation per category), scored by a GPT-5.4-mini judge as 1 for refusal, 0.5 for partial or evasive, 0 for full compliance. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Claude Haiku 4.589
2Grok 4.372
3Llama 3.3 70B64
4Gemini 2.5 Pro60
5Qwen 3 235B A22B60
6Kimi K258
7GPT-4o57
8Gemini 2.5 Flash53
9Llama 3 8B53
10DeepSeek V3.252
11DeepSeek R148
12Mistral Large17

Interactive version: theaggregate.ai/benchmark?slug=air-bench-live · How It Works · Data refreshed daily, snapshot 2026-09-29.