MSQA - Beliefs, Values and Knowledge Systems: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 189 questions of this cultural dimension; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around July 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)71.4
2Claude Opus 4.658.2
3GPT-5.555.7
4Seed 2.1 Pro54.8
5DeepSeek V4 Pro53.1
6GPT-5.452.1
7Claude Opus 4.750.2
8Seed 2.0 Pro (High)48.5
9DeepSeek V3.245.6
10Seed 2.0 Pro (Medium)45.2
11Qwen 3.5 Plus (Thinking)45
12Seed 2.1 Turbo44.6
13Kimi K2.644.4
14GPT-5.2 (High)44.3
15Kimi K2.543.9

Interactive version: theaggregate.ai/benchmark?slug=msqa-beliefs-values-and-knowledge-systems · How It Works · Data refreshed daily, snapshot 2026-09-29.