EuraGovExam - Administration: leaderboard

Metric: Accuracy (%) on the administration questions across the five regions; image-only: the model sees one scanned multiple-choice civil-service exam question in its original language and layout, with a fixed answer-format instruction and no OCR or tools; greedy decoding, one run; answers not in the required final-line format count as wrong; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1GPT-583.67#91
2Gemini 2.5 Pro83.57#145
3O381.53#121
4Gemini 3 Flash (Preview)79.16#78
5Gemini 3 Pro (Preview)77.23#64
6O4 Mini71.21#172
7Gemini 2.5 Flash68.53#237
8GPT-5.268.53#105
9Claude Sonnet 464.12#194
10GPT-5 Nano63.59#415
11GPT-4.154.57#240
12GPT-4.1 Mini54.14#346
13GPT-4o47.48#333
14Qwen 2 VL 7B Instruct41.78#816
15Gemini 2.5 Flash Lite37.27#413

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=euragovexam-administration · How It Works · Data refreshed daily, snapshot 2026-10-11.