ArabicDialectSafety - Harm Category Classification: leaderboard
Metric: Macro-F1 (%; which of seven harm categories a test prompt belongs to, six Arabic varieties; 14-shot prompting with two examples per category, greedy decoding with at most 8 new tokens, dialect-blind input). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Fanar-1-9B-Instruct | 69 |
| 2 | Llama 3.1 8B Instruct | 60 |
| 3 | Jais-2-8B-Chat | 60 |
| 4 | Qwen 2.5 7B Instruct | 57 |
Interactive version: theaggregate.ai/benchmark?slug=arabicdialectsafety-harm-category-classification · How It Works · Data refreshed daily, snapshot 2026-09-26.