MSQA - Japanese: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 83 Japanese-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around February 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)52.8
2GPT-5.548
3Claude Opus 4.742
4GPT-5.438.6
5Claude Opus 4.635.5
6DeepSeek V4 Pro34.9
7Seed 2.1 Pro31.8
8GPT-5.2 (High)30.1
9Seed 2.0 Pro (High)27.6
10Seed 2.0 Pro (Medium)25.9
11Gemini 2.5 Flash23.7
12Seed 2.0 Lite22.5
13DeepSeek V3.221.8
14Seed 2.1 Turbo21.8
15GLM-517.7

Interactive version: theaggregate.ai/benchmark?slug=msqa-japanese · How It Works · Data refreshed daily, snapshot 2026-09-29.