EuraGovExam - Politics: leaderboard

Metric: Accuracy (%) on the politics questions across the five regions; image-only: the model sees one scanned multiple-choice civil-service exam question in its original language and layout, with a fixed answer-format instruction and no OCR or tools; greedy decoding, one run; answers not in the required final-line format count as wrong; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1GPT-579.9#91
2O379.9#121
3Gemini 2.5 Pro79.43#145
4O4 Mini73.21#172
5GPT-5.267.94#105
6Gemini 3 Pro (Preview)64.11#64
7Claude Sonnet 463.16#194
8Gemini 3 Flash (Preview)62.68#78
9Gemini 2.5 Flash62.2#237
10GPT-5 Nano57.42#415
11GPT-4.1 Mini47.37#346
12GPT-4.145.93#240
13GPT-4o34.93#333
14Qwen 2 VL 7B Instruct29.67#816
15Qwen 2.5 VL 7B Instruct24.4#643

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=euragovexam-politics · How It Works · Data refreshed daily, snapshot 2026-10-11.