UA-JudgeExam: leaderboard

Metric: Full-condition accuracy (%; question and four options shown, one letter answered, on a 400-item sample of the 8,128 items that remain after the Claude Haiku 4.5 option-only gate; denominators exclude unparseable answers (DeepSeek R1 352, Pixtral Large 365, Nova Micro 388, Llama 3.1 8B 202); greedy decoding where the provider accepts a temperature; chance is 25). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol95.8
2Claude Sonnet 4.674
3DeepSeek R166.8
4Nova Pro58
5Claude Haiku 4.557.5
6Qwen 3 32B (Non-reasoning)55.2
7Llama 3.3 70B Instruct54
8Gemma 3 12B (IT)47.2
9Ministral-3-8B-Instruct-251246.3
10Nova 2.0 Lite (Non-reasoning)46
11Nova Micro43.6

Interactive version: theaggregate.ai/benchmark?slug=ua-judgeexam · How It Works · Data refreshed daily, snapshot 2026-09-29.