MSQA - English: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 151 English-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around January 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)74.9
2Claude Opus 4.662.1
3DeepSeek V4 Pro60.5
4GPT-5.558.8
5Claude Opus 4.758.4
6Seed 2.1 Pro57.8
7GPT-5.457.2
8Kimi K2.654.8
9Seed 2.0 Pro (High)53.3
10GLM-553
11Kimi K2.551.4
12GPT-5.2 (High)51
13DeepSeek V3.250.7
14Seed 2.1 Turbo50.3
15Seed 2.0 Pro (Medium)50.3

Interactive version: theaggregate.ai/benchmark?slug=msqa-english · How It Works · Data refreshed daily, snapshot 2026-09-29.