MSQA - Social Norms and Customs: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 186 questions of this cultural dimension; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around March 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)67
2GPT-5.554.3
3Claude Opus 4.651.6
4GPT-5.451.1
5Seed 2.1 Pro50.5
6DeepSeek V4 Pro49.5
7Claude Opus 4.748.5
8Seed 2.0 Pro (High)45.6
9DeepSeek V3.244.1
10Seed 2.1 Turbo43.4
11GPT-5.2 (High)43
12Seed 2.0 Pro (Medium)42
13Seed 2.0 Lite41.3
14GLM-540.8
15Gemini 2.5 Flash40.4

Interactive version: theaggregate.ai/benchmark?slug=msqa-social-norms-and-customs · How It Works · Data refreshed daily, snapshot 2026-09-29.