MSQA - Spanish: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 80 Spanish-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around April 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)68.2
2Claude Opus 4.663.9
3Claude Opus 4.761.6
4DeepSeek V4 Pro56.1
5Seed 2.1 Pro55.1
6GPT-5.554.8
7Kimi K2.554.3
8GPT-5.452.2
9Kimi K2.652.1
10Seed 2.0 Pro (High)52.1
11GLM-551.8
12GPT-5.2 (High)51.3
13Qwen 3.5 Plus (Thinking)51.2
14Seed 2.0 Pro (Medium)50.8
15DeepSeek V3.250.7

Interactive version: theaggregate.ai/benchmark?slug=msqa-spanish · How It Works · Data refreshed daily, snapshot 2026-09-29.