MSQA - Malay: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 82 Malay-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around March 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)70.4
2GPT-5.551.8
3GPT-5.443.6
4Seed 2.1 Pro40.3
5Claude Opus 4.639.6
6GPT-5.2 (High)39.3
7Claude Opus 4.739.1
8DeepSeek V4 Pro37.9
9Gemini 2.5 Flash35.7
10Kimi K2.635.2
11Kimi K2.535.1
12GLM-533.6
13Seed 2.0 Pro (High)33.4
14Seed 2.0 Pro (Medium)30.1
15Seed 2.1 Turbo27.4

Interactive version: theaggregate.ai/benchmark?slug=msqa-malay · How It Works · Data refreshed daily, snapshot 2026-09-29.