EuraGovExam - Computer Science: leaderboard

Metric: Accuracy (%) on the computer science questions across the five regions; image-only: the model sees one scanned multiple-choice civil-service exam question in its original language and layout, with a fixed answer-format instruction and no OCR or tools; greedy decoding, one run; answers not in the required final-line format count as wrong; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1O4 Mini92.52#172
2Gemini 2.5 Pro92.06#145
3GPT-592.06#91
4O390.93#121
5GPT-5.285.94#105
6Gemini 3 Flash (Preview)85.94#78
7GPT-5 Nano83.67#415
8GPT-4.1 Mini75.74#346
9Gemini 2.5 Flash75.28#237
10GPT-4.173.47#240
11Gemini 3 Pro (Preview)73.47#64
12Claude Sonnet 469.39#194
13GPT-4o58.96#333
14Qwen 2.5 VL 7B Instruct41.72#643
15Qwen 2 VL 7B Instruct41.04#816

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=euragovexam-computer-science · How It Works · Data refreshed daily, snapshot 2026-10-11.