Judge Arena — leaderboard
Atla AI evaluation of LLM-as-judge capability via human preference. 28 models rated by ELO from 4,689 pairwise votes measuring evaluation quality, consistency, and alignment with human judgments.
Metric: ELO Score. Source: huggingface.co. Status: saturation imminent. 28 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.3 70B Instruct | 1335 |
| 2 | Claude 3 Opus | 1312 |
| 3 | GPT-4o | 1308 |
| 4 | GPT-4 Turbo | 1304 |
| 5 | Claude 3 Haiku | 1286 |
| 6 | Claude 3.5 Haiku | 1282 |
| 7 | GPT-3.5 Turbo | 1271 |
| 8 | Qwen 2.5 72B Instruct | 1269 |
| 9 | Llama 3.1 405B Instruct | 1263 |
| 10 | Llama 3.1 8B Instruct | 1233 |
| 11 | Mistral 7B Instruct (v0.3) | 1215 |
| 12 | Claude 3.5 Sonnet | 1211 |
| 13 | QwQ 32B-Preview | 1172 |
| 14 | Mistral 7B Instruct (v0.1) | 1157 |
| 15 | Qwen 2 72B Instruct | 1123 |
Interactive version: theaggregate.ai/benchmark?slug=judge-arena · How the rankings work · Data refreshed daily, snapshot 2026-07-22.