CultureTalk-ID - Local Language Dialogue MCQ: leaderboard
Metric: Accuracy (%; the same dialogues in their local languages (Javanese dialects, Acehnese, Minangkabau, Balinese, Banjarese, Buginese, Uab Meto and Wamesa); three-option choice of the culturally appropriate final utterance of a dialogue grounded in a regional Indonesian culture, no province or language context in the prompt, averaged over province-specific and general items (test split); higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 87.77 |
| 2 | GPT-5.1 | 82.31 |
| 3 | Command A | 75.09 |
| 4 | Sailor2-8B-Chat | 70.15 |
| 5 | gemma2-9B-cpt-sahabatai-v1-instruct | 69.08 |
| 6 | Gemma 2 9B (IT) | 65.17 |
| 7 | Qwen 3 8B | 52.89 |
| 8 | Llama 3.1 8B Instruct | 51.87 |
Interactive version: theaggregate.ai/benchmark?slug=culturetalk-id-local-language-dialogue-mcq · How It Works · Data refreshed daily, snapshot 2026-09-29.