CultureTalk-ID - Indonesian Dialogue MCQ: leaderboard

Metric: Accuracy (%; Indonesian dialogues; three-option choice of the culturally appropriate final utterance of a dialogue grounded in a regional Indonesian culture, no province or language context in the prompt, averaged over province-specific and general items (test split); higher is better). Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash91.46
2GPT-5.189.23
3Command A84.74
4gemma2-9B-cpt-sahabatai-v1-instruct80.87
5Sailor2-8B-Chat80.07
6Gemma 2 9B (IT)76.78
7Qwen 3 8B68.95
8Llama 3.1 8B Instruct62.63

Interactive version: theaggregate.ai/benchmark?slug=culturetalk-id-indonesian-dialogue-mcq · How It Works · Data refreshed daily, snapshot 2026-09-29.