EuraGovExam - Mathematics: leaderboard

Metric: Accuracy (%) on the mathematics questions across the five regions; image-only: the model sees one scanned multiple-choice civil-service exam question in its original language and layout, with a fixed answer-format instruction and no OCR or tools; greedy decoding, one run; answers not in the required final-line format count as wrong; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro86.28#145
2O4 Mini80.61#172
3O378.98#121
4GPT-578.22#91
5Gemini 3 Flash (Preview)73.22#78
6Gemini 2.5 Flash73.13#237
7GPT-5 Nano72.46#415
8GPT-5.268.52#105
9Gemini 3 Pro (Preview)64.88#64
10Claude Sonnet 461.61#194
11GPT-4.158.06#240
12GPT-4.1 Mini56.24#346
13GPT-4o43.38#333
14Qwen 2.5 VL 7B Instruct28.02#643
15Qwen 2 VL 7B Instruct20.73#816

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=euragovexam-mathematics · How It Works · Data refreshed daily, snapshot 2026-10-11.