MSQA - Cultural Products and Symbols: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 208 questions of this cultural dimension; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around May 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)66.1
2GPT-5.555
3Claude Opus 4.650.6
4GPT-5.449.2
5GPT-5.2 (High)46.7
6DeepSeek V4 Pro46.2
7Claude Opus 4.745.6
8Seed 2.1 Pro45.5
9Qwen 3.5 Plus (Thinking)41.4
10GLM-539.8
11Kimi K2.639.3
12Seed 2.0 Pro (High)38.9
13Kimi K2.538.7
14DeepSeek V3.238.7
15Seed 2.0 Pro (Medium)38.4

Interactive version: theaggregate.ai/benchmark?slug=msqa-cultural-products-and-symbols · How It Works · Data refreshed daily, snapshot 2026-09-29.