ArabCulture-Dialogue - Dialect Steering Quality: leaderboard

Metric: Continuation quality (0-100): GPT-5 judge rating on 1-5 rescaled as (s-1)/4 and reported times 100, zero-shot dialect steering: given a dialogue context and an MSA utterance, the model writes one continuation in the target country's dialect, over the dialogues of ArabCulture-Dialogue (13 Arab countries, parallel Modern Standard Arabic and country-dialect versions of 3,471 culturally grounded multi-turn dialogues built from ArabCulture); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScore
1GPT-595.6
2Gemini 2.5 Pro92.4
3ALLaM-7B-Instruct-preview81.68
4Jais-2-8B-Chat73
5c4ai-command-r7B-arabic-02-202571.6
6Fanar-1-9B62.1
7Hala-9B61.2
8Gemma 2 9B (IT)60.4
9jais-adapted-7B-chat59.9
10SILMA-9B-Instruct-v1.059.1
11Llama 3.1 8B Instruct51.6
12Qwen 3 8B43

Interactive version: theaggregate.ai/benchmark?slug=arabculture-dialogue-dialect-steering-quality · How It Works · Data refreshed daily, snapshot 2026-10-07.