SalamahBench (MSA) - Macro ASR: leaderboard

Metric: Macro attack success rate (%) on the Modern Standard Arabic (MSA) half of SalamahBench (8,270 human-verified harmful prompts): the unweighted mean of the per-category attack success rates over the eleven named ML Commons hazard categories; Qwen3Guard (generative, strict setting: responses labelled Unsafe or Controversial count as attack successes) judges each response; temperature 0; lower is better. Source: arxiv.org. 6 models tracked.

Top models

#ModelScoreOverall rank
1Falcon-H1R-7B4.5#658
2Jais-2-8B-Chat26.1#838

Interactive version: theaggregate.ai/benchmark?slug=salamahbench-msa-macro-asr · How It Works · Data refreshed daily, snapshot 2026-10-11.