EuraGovExam - Law: leaderboard

Metric: Accuracy (%) on the law questions across the five regions; image-only: the model sees one scanned multiple-choice civil-service exam question in its original language and layout, with a fixed answer-format instruction and no OCR or tools; greedy decoding, one run; answers not in the required final-line format count as wrong; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 28 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro84.24#145
2GPT-581.03#91
3O378.08#121
4Gemini 3 Flash (Preview)76.48#78
5Gemini 3 Pro (Preview)71.67#64
6GPT-5.263.92#105
7O4 Mini63.92#172
8Gemini 2.5 Flash63.18#237
9Claude Sonnet 460.84#194
10GPT-5 Nano57.14#415
11GPT-4.1 Mini51.23#346
12GPT-4.147.54#240
13GPT-4o42#333
14Qwen 2 VL 7B Instruct37.81#816
15Qwen 2.5 VL 7B Instruct34.48#643

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=euragovexam-law · How It Works · Data refreshed daily, snapshot 2026-10-11.