CultureTalk-ID - Translation into Local Languages: leaderboard
Metric: BLEU-4 (0-100) of zero-shot dialogue translation from Indonesian into the local languages with province and language context, mean of three runs for open-weight models; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 28.36 |
| 2 | GPT-5.1 | 27.53 |
| 3 | gemma2-9B-cpt-sahabatai-v1-instruct | 21.7 |
| 4 | Command A | 21.45 |
| 5 | Gemma 2 9B (IT) | 20.54 |
| 6 | Qwen 3 8B | 19.78 |
| 7 | Llama 3.1 8B Instruct | 17.62 |
| 8 | Sailor2-8B-Chat | 15.44 |
Interactive version: theaggregate.ai/benchmark?slug=culturetalk-id-translation-into-local-languages · How It Works · Data refreshed daily, snapshot 2026-09-29.