ArabicDialectSafety - Harm Category Classification: leaderboard

Metric: Macro-F1 (%; which of seven harm categories a test prompt belongs to, six Arabic varieties; 14-shot prompting with two examples per category, greedy decoding with at most 8 new tokens, dialect-blind input). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1Fanar-1-9B-Instruct69
2Llama 3.1 8B Instruct60
3Jais-2-8B-Chat60
4Qwen 2.5 7B Instruct57

Interactive version: theaggregate.ai/benchmark?slug=arabicdialectsafety-harm-category-classification · How It Works · Data refreshed daily, snapshot 2026-09-26.