ProLLM - LLM-as-a-Judge — leaderboard

ProLLM benchmark evaluating how well LLMs can judge and score the outputs of other AI models - measuring meta-evaluation capability.

Metric: Score (%). Source: www.prollm.ai. Status: saturated. 58 models tracked.

Top models

#ModelScore
1GPT-4.185.3
2GPT-4o84.6
3GPT-OSS-120B84.6
4GPT-4 Turbo83.8
5O1 Mini83.8
6DeepSeek V382.4
7O182.4
8GPT-4.582.4
9O3 Mini (Medium)82.4
10O3 Mini (High)81.6
11Qwen 2.5 72B Instruct80.9
12Llama 4 Maverick Instruct80.9
13GPT-4.1 Mini80.1
14Llama 3.3 70B Instruct79.4
15Gemma 2 27B (IT)79.4

Interactive version: theaggregate.ai/benchmark?slug=prollm-llm-as-a-judge · How the rankings work · Data refreshed daily, snapshot 2026-07-22.