EuraGovExam - South Korea: leaderboard

Metric: Accuracy (%) on the 2,445 questions from South Korea (national civil service papers); image-only: the model sees one scanned multiple-choice civil-service exam question in its original language and layout, with a fixed answer-format instruction and no OCR or tools; greedy decoding, one run; answers not in the required final-line format count as wrong; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1GPT-591.17#91
2Gemini 2.5 Pro91.12#145
3O390.06#121
4Gemini 3 Flash (Preview)84.5#78
5O4 Mini82.49#172
6Gemini 3 Pro (Preview)75.26#64
7GPT-5.273.09#105
8GPT-5 Nano72.97#415
9Gemini 2.5 Flash67.65#237
10Claude Sonnet 462.41#194
11GPT-4.1 Mini59.92#346
12GPT-4.154.23#240
13GPT-4o33.25#333
14Qwen 2.5 VL 7B Instruct29.53#643
15Qwen 2 VL 7B Instruct29.08#816

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=euragovexam-south-korea · How It Works · Data refreshed daily, snapshot 2026-10-11.