SalamahBench: leaderboard
Metric: Attack success rate (%) over all 49,620 SalamahBench instances (8,270 human-verified harmful prompts in Modern Standard Arabic and five regional Arabic varieties: Egyptian, Syrian, Saudi, Lebanese and Moroccan); Qwen3Guard (generative, strict setting: responses labelled Unsafe or Controversial count as attack successes) judges each response; temperature 0; lower is better. Source: arxiv.org. 3 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | ALLaM 2 7B | 4.4 | |
| 2 | Fanar-2-27B-Instruct | 9.2 | |
| 3 | Karnak | 19.8 |
Interactive version: theaggregate.ai/benchmark?slug=salamahbench · How It Works · Data refreshed daily, snapshot 2026-10-11.