MSQA - Korean: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 86 Korean-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around December 2026. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)73.9
2Claude Opus 4.762.8
3GPT-5.562.4
4GPT-5.459.7
5Claude Opus 4.656.4
6GPT-5.2 (High)46.6
7DeepSeek V4 Pro45.8
8Seed 2.1 Pro45.2
9Gemini 2.5 Flash41.2
10Seed 2.0 Pro (High)34
11Seed 2.0 Pro (Medium)32.6
12Seed 2.1 Turbo32
13Qwen 3.5 Plus (Thinking)31.7
14GLM-529.3
15Seed 2.0 Lite28.7

Interactive version: theaggregate.ai/benchmark?slug=msqa-korean · How It Works · Data refreshed daily, snapshot 2026-09-29.