SalamahBench: leaderboard

Metric: Attack success rate (%) over all 49,620 SalamahBench instances (8,270 human-verified harmful prompts in Modern Standard Arabic and five regional Arabic varieties: Egyptian, Syrian, Saudi, Lebanese and Moroccan); Qwen3Guard (generative, strict setting: responses labelled Unsafe or Controversial count as attack successes) judges each response; temperature 0; lower is better. Source: arxiv.org. 3 models tracked.

Top models

#ModelScoreOverall rank
1ALLaM 2 7B4.4
2Fanar-2-27B-Instruct9.2
3Karnak19.8

Interactive version: theaggregate.ai/benchmark?slug=salamahbench · How It Works · Data refreshed daily, snapshot 2026-10-11.