SLMJury — leaderboard
Evaluates small language models (0.6B-14B) as judges: each model scores responses and its verdicts are compared against ground-truth labels, reporting the best judge accuracy per model across thinking-token-budget configurations.
Metric: Judge accuracy (%). Source: anishh15.github.io. Status: saturated. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Phi-4 | 89.67 |
| 2 | Qwen 3 14B | 89.51 |
| 3 | Qwen 3 4B | 89.2 |
| 4 | Qwen 3 8B | 88.96 |
| 5 | Phi-4 (Reasoning) | 88.24 |
| 6 | Phi-4 Mini Instruct | 88.22 |
| 7 | Llama 3.1 8B Instruct | 86.79 |
| 8 | Qwen 2.5 7B Instruct | 86.62 |
| 9 | Qwen 3 1.7B | 85.96 |
| 10 | Phi-4-mini (Reasoning) | 83.32 |
| 11 | Llama 3.2 3B Instruct | 81.77 |
| 12 | Qwen 2.5 1.5B Instruct | 79.91 |
| 13 | Qwen 2.5 3B Instruct | 77.65 |
| 14 | Qwen 3 0.6B | 76.37 |
| 15 | Llama 3.2 1B Instruct | 69.38 |
Interactive version: theaggregate.ai/benchmark?slug=slmjury · How the rankings work · Data refreshed daily, snapshot 2026-07-22.