ArabCulture-Dialogue - Dialect Steering Quality: leaderboard
Metric: Continuation quality (0-100): GPT-5 judge rating on 1-5 rescaled as (s-1)/4 and reported times 100, zero-shot dialect steering: given a dialogue context and an MSA utterance, the model writes one continuation in the target country's dialect, over the dialogues of ArabCulture-Dialogue (13 Arab countries, parallel Modern Standard Arabic and country-dialect versions of 3,471 culturally grounded multi-turn dialogues built from ArabCulture); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 95.6 |
| 2 | Gemini 2.5 Pro | 92.4 |
| 3 | ALLaM-7B-Instruct-preview | 81.68 |
| 4 | Jais-2-8B-Chat | 73 |
| 5 | c4ai-command-r7B-arabic-02-2025 | 71.6 |
| 6 | Fanar-1-9B | 62.1 |
| 7 | Hala-9B | 61.2 |
| 8 | Gemma 2 9B (IT) | 60.4 |
| 9 | jais-adapted-7B-chat | 59.9 |
| 10 | SILMA-9B-Instruct-v1.0 | 59.1 |
| 11 | Llama 3.1 8B Instruct | 51.6 |
| 12 | Qwen 3 8B | 43 |
Interactive version: theaggregate.ai/benchmark?slug=arabculture-dialogue-dialect-steering-quality · How It Works · Data refreshed daily, snapshot 2026-10-07.