SALAD-Bench — leaderboard
Hierarchical safety benchmark for LLMs, attacks, and defenses, with a three-level taxonomy and variants for standard, attack-enhanced, defense-enhanced, and multiple-choice safety evaluation.
Metric: Average Safety Score (%). Source: huggingface.co. Status: saturation imminent. 34 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 2 | 93.91 |
| 2 | GPT-4 Turbo | 86.88 |
| 3 | Llama 2 13B Chat Base | 81.27 |
| 4 | Llama 2 70B Chat (HF) | 81.22 |
| 5 | GPT-3.5 Turbo | 81.01 |
| 6 | gemma-2B (IT) | 73.12 |
| 7 | internlm-20B Chat | 63.31 |
| 8 | Qwen 1.5 4B Chat | 59.43 |
| 9 | internlm2-chat-7B | 58.99 |
| 10 | Llama 2 7B Chat | 57.38 |
| 11 | Yi 34B (Chat) | 55.38 |
| 12 | Qwen 1.5 72B Chat | 54.88 |
| 13 | gemma-7B (IT) | 54.81 |
| 14 | internlm2-chat-20B | 54.61 |
| 15 | Gemini 1.0 Pro | 54.17 |
Interactive version: theaggregate.ai/benchmark?slug=salad-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.