MSQA - Indonesian: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 81 Indonesian-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around January 2028. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)62.9
2Claude Opus 4.761.3
3Claude Opus 4.648.5
4GPT-5.545.4
5Seed 2.1 Pro45.2
6GPT-5.445
7DeepSeek V4 Pro43.1
8DeepSeek V3.242.8
9Seed 2.0 Pro (High)42.6
10Gemini 2.5 Flash40.6
11Seed 2.1 Turbo40
12Qwen 3.5 Plus (Thinking)38.9
13GPT-5.2 (High)37.3
14Seed 2.0 Pro (Medium)37.3
15Kimi K2.636.7

Interactive version: theaggregate.ai/benchmark?slug=msqa-indonesian · How It Works · Data refreshed daily, snapshot 2026-09-29.