Judge Arena — leaderboard

Atla AI evaluation of LLM-as-judge capability via human preference. 28 models rated by ELO from 4,689 pairwise votes measuring evaluation quality, consistency, and alignment with human judgments.

Metric: ELO Score. Source: huggingface.co. Status: saturation imminent. 28 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct1335
2Claude 3 Opus1312
3GPT-4o1308
4GPT-4 Turbo1304
5Claude 3 Haiku1286
6Claude 3.5 Haiku1282
7GPT-3.5 Turbo1271
8Qwen 2.5 72B Instruct1269
9Llama 3.1 405B Instruct1263
10Llama 3.1 8B Instruct1233
11Mistral 7B Instruct (v0.3)1215
12Claude 3.5 Sonnet1211
13QwQ 32B-Preview1172
14Mistral 7B Instruct (v0.1)1157
15Qwen 2 72B Instruct1123

Interactive version: theaggregate.ai/benchmark?slug=judge-arena · How the rankings work · Data refreshed daily, snapshot 2026-07-22.