MSQA - Language Expression and Communication Arts: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 220 questions of this cultural dimension; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around December 2026. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)70.7
2GPT-5.560.7
3GPT-5.457.3
4Claude Opus 4.655.9
5Seed 2.1 Pro54.6
6DeepSeek V4 Pro54.1
7GPT-5.2 (High)52.5
8Claude Opus 4.749.2
9GLM-543.8
10Seed 2.0 Pro (High)43.6
11Gemini 2.5 Flash43.5
12Qwen 3.5 Plus (Thinking)41.8
13Seed 2.0 Pro (Medium)41.3
14Kimi K2.640.9
15Kimi K2.539.9

Interactive version: theaggregate.ai/benchmark?slug=msqa-language-expression-and-communication-arts · How It Works · Data refreshed daily, snapshot 2026-09-29.