Judge Arena: leaderboard

Atla AI evaluation of LLM-as-judge capability via human preference. 28 models rated by ELO from 4,689 pairwise votes measuring evaluation quality, consistency, and alignment with human judgments.

Metric: ELO Score. Source: huggingface.co. Status: saturated. 28 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct1335
2Claude 3 Opus1312
3GPT-4o1308
4GPT-4 Turbo1304
5Claude 3 Haiku1286
6Claude 3.5 Haiku1282
7Qwen 2.5 7B Instruct Turbo1280
8GPT-3.5 Turbo1271
9Qwen 2.5 72B Instruct Turbo1269
10Llama 3.1 405B Instruct1263
11Llama 3.1 8B Instruct1233
12Mistral 7B Instruct (v0.3)1215
13Claude 3.5 Sonnet1211
14QwQ 32B-Preview1172
15Mistral 7B Instruct (v0.1)1157

Interactive version: theaggregate.ai/benchmark?slug=judge-arena · How It Works · Data refreshed daily, snapshot 2026-09-05.