EuraGovExam - Medicine: leaderboard

Metric: Accuracy (%) on the medicine questions across the five regions; image-only: the model sees one scanned multiple-choice civil-service exam question in its original language and layout, with a fixed answer-format instruction and no OCR or tools; greedy decoding, one run; answers not in the required final-line format count as wrong; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro95.8#145
2GPT-593.28#91
3Gemini 3 Flash (Preview)92.86#78
4O391.18#121
5Gemini 3 Pro (Preview)88.66#64
6O4 Mini88.24#172
7Gemini 2.5 Flash83.19#237
8GPT-5 Nano81.51#415
9GPT-5.281.09#105
10Claude Sonnet 478.99#194
11GPT-4.1 Mini72.69#346
12GPT-4.169.75#240
13GPT-4o50#333
14Gemini 2.5 Flash Lite42.02#413
15Qwen 2 VL 7B Instruct37.82#816

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=euragovexam-medicine · How It Works · Data refreshed daily, snapshot 2026-10-11.