Judgemark v2.1 — leaderboard

Evaluates LLM-as-judge quality: stability (self-consistency), separability (ability to distinguish model quality), and correlation with human preferences.

Metric: Judgemark Score (0-100). Source: eqbench.com. Status: saturation imminent. 47 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.694.08
2Claude Opus 4.692.93
3Gemini 3 Pro (Preview)86.58
4GLM-585.58
5Qwen 3.5 397B A17B85.45
6Gemini 3 Flash (Preview)85.38
7GPT-5.484.98
8Claude Sonnet 4.584.56
9Qwen 3.5 27B84.31
10GPT-5.282.01
11Claude Sonnet 481.99
12Qwen 3.5 35B A3B81.39
13GPT-5.4 Mini81.14
14Qwen 3.5 122B A10B81.03
15Gemini 3.1 Flash Lite (Preview)80.92

Interactive version: theaggregate.ai/benchmark?slug=judgemark-v2-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.