MSQA - Thai: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 95 Thai-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around May 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)71
2GPT-5.558.5
3Claude Opus 4.651
4DeepSeek V4 Pro49.5
5GPT-5.448.6
6Seed 2.1 Pro45.9
7GPT-5.2 (High)45.1
8Gemini 2.5 Flash40.2
9Qwen 3.5 Plus (Thinking)39.3
10DeepSeek V3.234.8
11GLM-534.8
12Seed 2.0 Lite34.3
13Seed 2.0 Pro (High)30.9
14Kimi K2.628.3
15Kimi K2.527.9

Interactive version: theaggregate.ai/benchmark?slug=msqa-thai · How It Works · Data refreshed daily, snapshot 2026-09-29.