SLMJury — leaderboard

Evaluates small language models (0.6B-14B) as judges: each model scores responses and its verdicts are compared against ground-truth labels, reporting the best judge accuracy per model across thinking-token-budget configurations.

Metric: Judge accuracy (%). Source: anishh15.github.io. Status: saturated. 16 models tracked.

Top models

#ModelScore
1Phi-489.67
2Qwen 3 14B89.51
3Qwen 3 4B89.2
4Qwen 3 8B88.96
5Phi-4 (Reasoning)88.24
6Phi-4 Mini Instruct88.22
7Llama 3.1 8B Instruct86.79
8Qwen 2.5 7B Instruct86.62
9Qwen 3 1.7B85.96
10Phi-4-mini (Reasoning)83.32
11Llama 3.2 3B Instruct81.77
12Qwen 2.5 1.5B Instruct79.91
13Qwen 2.5 3B Instruct77.65
14Qwen 3 0.6B76.37
15Llama 3.2 1B Instruct69.38

Interactive version: theaggregate.ai/benchmark?slug=slmjury · How the rankings work · Data refreshed daily, snapshot 2026-07-22.