ArabicDialectSafety - Unsafe Response Rate: leaderboard

Metric: Unsafe response rate (%; share of 500 stratified harmful dialectal Arabic test prompts whose response a GPT-5.2 judge labels unsafe under a binary rubric; lower is better). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Qwen 3.6 Plus0.2
2Claude Sonnet 4.60.6
3Claude 3 Haiku1.6
4Gemini 2.5 Flash2
5GPT-4o Mini4.6

Interactive version: theaggregate.ai/benchmark?slug=arabicdialectsafety-unsafe-response-rate · How It Works · Data refreshed daily, snapshot 2026-09-26.