MSQA - Portuguese: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 80 Portuguese-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around March 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)70.7
2Claude Opus 4.767.3
3Seed 2.1 Pro61.8
4GPT-5.559.4
5DeepSeek V4 Pro59
6Gemini 2.5 Flash58.6
7GPT-5.2 (High)57.5
8GPT-5.457.3
9Seed 2.0 Pro (High)56.1
10DeepSeek V3.255.7
11Qwen 3.5 Plus (Thinking)54.4
12GLM-554.1
13Seed 2.0 Pro (Medium)53.1
14Seed 2.1 Turbo52.4
15Kimi K2.552

Interactive version: theaggregate.ai/benchmark?slug=msqa-portuguese · How It Works · Data refreshed daily, snapshot 2026-09-29.