MSQA - Chinese: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 150 Chinese-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around 2029. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)56.8
2Seed 2.0 Pro (High)55.7
3DeepSeek V4 Pro50.9
4Seed 2.1 Pro50.7
5Seed 2.0 Pro (Medium)50.2
6Seed 2.1 Turbo48.7
7DeepSeek V3.245.3
8GPT-5.544.5
9Claude Opus 4.643.9
10GLM-543.5
11Seed 2.0 Lite42.2
12Kimi K2.541.5
13Kimi K2.640.9
14GPT-5.438.3
15Qwen 3.5 Plus (Thinking)34.9

Interactive version: theaggregate.ai/benchmark?slug=msqa-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.