ArabCulture-Dialogue - Dialect Steering: leaderboard
Metric: Dialect identity accuracy (%): share of continuations whose GlotLID language code exactly matches the target country's dialect code (strict ISO 639-3), zero-shot dialect steering: given a dialogue context and an MSA utterance, the model writes one continuation in the target country's dialect, over the dialogues of ArabCulture-Dialogue (13 Arab countries, parallel Modern Standard Arabic and country-dialect versions of 3,471 culturally grounded multi-turn dialogues built from ArabCulture); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 50.5 |
| 2 | GPT-5 | 45.4 |
| 3 | ALLaM-7B-Instruct-preview | 36.25 |
| 4 | Jais-2-8B-Chat | 20.8 |
| 5 | c4ai-command-r7B-arabic-02-2025 | 18.6 |
| 6 | Gemma 2 9B (IT) | 16.4 |
| 7 | SILMA-9B-Instruct-v1.0 | 7.8 |
| 8 | Llama 3.1 8B Instruct | 6.4 |
| 9 | Qwen 3 8B | 4.1 |
| 10 | Fanar-1-9B | 3.8 |
| 11 | jais-adapted-7B-chat | 2.2 |
| 12 | Hala-9B | 1.6 |
Interactive version: theaggregate.ai/benchmark?slug=arabculture-dialogue-dialect-steering · How It Works · Data refreshed daily, snapshot 2026-10-07.