MSQA - Russian: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 92 Russian-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around December 2026. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)77.1
2GPT-5.565.3
3GPT-5.462.6
4Claude Opus 4.659
5DeepSeek V4 Pro57.5
6Seed 2.1 Pro57.5
7GPT-5.2 (High)55.3
8Qwen 3.5 Plus (Thinking)48.3
9GLM-547
10Gemini 2.5 Flash46.3
11Seed 2.1 Turbo41.6
12Seed 2.0 Pro (High)39.8
13Kimi K2.638.2
14DeepSeek V3.237.7
15Seed 2.0 Lite37.3

Interactive version: theaggregate.ai/benchmark?slug=msqa-russian · How It Works · Data refreshed daily, snapshot 2026-09-29.