EuraGovExam: leaderboard

Metric: Accuracy (%) over all 8,000 questions from five regions and 17 domains; image-only: the model sees one scanned multiple-choice civil-service exam question in its original language and layout, with a fixed answer-format instruction and no OCR or tools; greedy decoding, one run; answers not in the required final-line format count as wrong; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro86.99#145
2GPT-585.8#91
3O384.26#121
4O4 Mini79.4#172
5Gemini 3 Flash (Preview)75.28#78
6GPT-5.269.94#105
7Gemini 3 Pro (Preview)68.53#64
8Gemini 2.5 Flash68.33#237
9GPT-5 Nano67.58#415
10Claude Sonnet 463.29#194
11GPT-4.1 Mini56.27#346
12GPT-4.154.73#240
13GPT-4o42.04#333
14Qwen 2.5 VL 7B Instruct32.3#643
15Qwen 2 VL 7B Instruct31.38#816

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=euragovexam · How It Works · Data refreshed daily, snapshot 2026-10-11.