CultureTalk-ID - Local Language Dialogue MCQ: leaderboard

Metric: Accuracy (%; the same dialogues in their local languages (Javanese dialects, Acehnese, Minangkabau, Balinese, Banjarese, Buginese, Uab Meto and Wamesa); three-option choice of the culturally appropriate final utterance of a dialogue grounded in a regional Indonesian culture, no province or language context in the prompt, averaged over province-specific and general items (test split); higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash87.77
2GPT-5.182.31
3Command A75.09
4Sailor2-8B-Chat70.15
5gemma2-9B-cpt-sahabatai-v1-instruct69.08
6Gemma 2 9B (IT)65.17
7Qwen 3 8B52.89
8Llama 3.1 8B Instruct51.87

Interactive version: theaggregate.ai/benchmark?slug=culturetalk-id-local-language-dialogue-mcq · How It Works · Data refreshed daily, snapshot 2026-09-29.