MSQA: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over all 1,064 questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around April 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)68.7
2GPT-5.555.6
3Claude Opus 4.652.8
4GPT-5.450.9
5DeepSeek V4 Pro50.8
6Seed 2.1 Pro50.4
7Claude Opus 4.749
8GPT-5.2 (High)45.4
9Seed 2.0 Pro (High)44.3
10GLM-542.1
11Seed 2.0 Pro (Medium)41.2
12Seed 2.1 Turbo40.8
13Gemini 2.5 Flash40.6
14DeepSeek V3.240.4
15Qwen 3.5 Plus (Thinking)40.4

Interactive version: theaggregate.ai/benchmark?slug=msqa · How It Works · Data refreshed daily, snapshot 2026-09-29.