ArabCulture-Dialogue - Dialect Steering: leaderboard

Metric: Dialect identity accuracy (%): share of continuations whose GlotLID language code exactly matches the target country's dialect code (strict ISO 639-3), zero-shot dialect steering: given a dialogue context and an MSA utterance, the model writes one continuation in the target country's dialect, over the dialogues of ArabCulture-Dialogue (13 Arab countries, parallel Modern Standard Arabic and country-dialect versions of 3,471 culturally grounded multi-turn dialogues built from ArabCulture); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro50.5
2GPT-545.4
3ALLaM-7B-Instruct-preview36.25
4Jais-2-8B-Chat20.8
5c4ai-command-r7B-arabic-02-202518.6
6Gemma 2 9B (IT)16.4
7SILMA-9B-Instruct-v1.07.8
8Llama 3.1 8B Instruct6.4
9Qwen 3 8B4.1
10Fanar-1-9B3.8
11jais-adapted-7B-chat2.2
12Hala-9B1.6

Interactive version: theaggregate.ai/benchmark?slug=arabculture-dialogue-dialect-steering · How It Works · Data refreshed daily, snapshot 2026-10-07.