CultureTalk-ID - Indonesian Dialogue MCQ: leaderboard
Metric: Accuracy (%; Indonesian dialogues; three-option choice of the culturally appropriate final utterance of a dialogue grounded in a regional Indonesian culture, no province or language context in the prompt, averaged over province-specific and general items (test split); higher is better). Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 91.46 |
| 2 | GPT-5.1 | 89.23 |
| 3 | Command A | 84.74 |
| 4 | gemma2-9B-cpt-sahabatai-v1-instruct | 80.87 |
| 5 | Sailor2-8B-Chat | 80.07 |
| 6 | Gemma 2 9B (IT) | 76.78 |
| 7 | Qwen 3 8B | 68.95 |
| 8 | Llama 3.1 8B Instruct | 62.63 |
Interactive version: theaggregate.ai/benchmark?slug=culturetalk-id-indonesian-dialogue-mcq · How It Works · Data refreshed daily, snapshot 2026-09-29.