EuraGovExam - Language: leaderboard

Metric: Accuracy (%) on the language questions across the five regions; image-only: the model sees one scanned multiple-choice civil-service exam question in its original language and layout, with a fixed answer-format instruction and no OCR or tools; greedy decoding, one run; answers not in the required final-line format count as wrong; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro91.49#145
2O391.49#121
3GPT-590.27#91
4Gemini 3 Pro (Preview)86.89#64
5Gemini 3 Flash (Preview)85.81#78
6O4 Mini84.19#172
7Gemini 2.5 Flash75.68#237
8Claude Sonnet 472.3#194
9GPT-4.1 Mini61.62#346
10GPT-5.261.35#105
11GPT-5 Nano61.08#415
12GPT-4.160.14#240
13GPT-4o49.73#333
14Qwen 2.5 VL 7B Instruct43.51#643
15Qwen 2 VL 7B Instruct41.89#816

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=euragovexam-language · How It Works · Data refreshed daily, snapshot 2026-10-11.