SALAD-Bench — leaderboard

Hierarchical safety benchmark for LLMs, attacks, and defenses, with a three-level taxonomy and variants for standard, attack-enhanced, defense-enhanced, and multiple-choice safety evaluation.

Metric: Average Safety Score (%). Source: huggingface.co. Status: saturation imminent. 34 models tracked.

Top models

#ModelScore
1Claude 293.91
2GPT-4 Turbo86.88
3Llama 2 13B Chat Base81.27
4Llama 2 70B Chat (HF)81.22
5GPT-3.5 Turbo81.01
6gemma-2B (IT)73.12
7internlm-20B Chat63.31
8Qwen 1.5 4B Chat59.43
9internlm2-chat-7B58.99
10Llama 2 7B Chat57.38
11Yi 34B (Chat)55.38
12Qwen 1.5 72B Chat54.88
13gemma-7B (IT)54.81
14internlm2-chat-20B54.61
15Gemini 1.0 Pro54.17

Interactive version: theaggregate.ai/benchmark?slug=salad-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.