MSQA - History and Collective Memory: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 261 questions of this cultural dimension; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around May 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)68.5
2GPT-5.552.5
3Claude Opus 4.751
4DeepSeek V4 Pro50.8
5Claude Opus 4.648.7
6Seed 2.1 Pro47.4
7GPT-5.445.9
8Seed 2.0 Pro (High)45.1
9GLM-542.2
10GPT-5.2 (High)40.7
11Seed 2.1 Turbo39.9
12Seed 2.0 Pro (Medium)39.7
13Seed 2.0 Lite39.3
14Kimi K2.638.4
15Gemini 2.5 Flash38.2

Interactive version: theaggregate.ai/benchmark?slug=msqa-history-and-collective-memory · How It Works · Data refreshed daily, snapshot 2026-09-29.