MSQA - French: leaderboard

Metric: F-score (%; harmonic mean of CO, the share of fully correct answers, and CGA, the correct share of attempted answers, over the 84 French-language questions; free-form answers to natively sourced culture-specific short questions asked in their own language, five runs per model, scored by a Gemini-3.1-Pro gold-answer containment judge). Source: arxiv.org. Saturation forecast: Around December 2026. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)79.7
2Claude Opus 4.768.9
3Claude Opus 4.667.8
4GPT-5.566.5
5GPT-5.460.8
6GPT-5.2 (High)58.5
7Seed 2.1 Pro57
8DeepSeek V4 Pro55.1
9Qwen 3.5 Plus (Thinking)53.2
10GLM-552.2
11Kimi K2.651.8
12Kimi K2.550.4
13Seed 2.0 Pro (Medium)48.2
14Seed 2.1 Turbo48.1
15Seed 2.0 Pro (High)48

Interactive version: theaggregate.ai/benchmark?slug=msqa-french · How It Works · Data refreshed daily, snapshot 2026-09-29.