RankJudge - Machine Learning: leaderboard

Metric: Bradley-Terry Elo rating (open scale, mean strength mapped to 1500) of an LLM judge, fitted on joint correctness (the judge must pick the flawed conversation, its flawed turn and its failure type) over RankJudge pairs of multi-turn, document-grounded conversations in which exactly one conversation carries one injected assistant failure; pairs every judge answers alike and the top 5 percent hardest pairs are dropped; machine-learning domain (240 pairs grounded in CS papers); higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)1992
2Claude Sonnet 4.61755
3Kimi K2.61743
4Gemma 4 31B1708
5Gemini 3 Flash1665
6Qwen 3.6 Plus1665
7GLM-5.11665
8Qwen 3.5 397B A17B1590
9GPT-5.41581
10Claude Opus 4.71523
11Claude Haiku 4.51492
12Gemma 4 26B A4B1428
13Qwen 3.5 122B A10B1381
14Qwen 3.5 35B A3B1349
15GPT-5.4 Mini1343

Interactive version: theaggregate.ai/benchmark?slug=rankjudge-machine-learning · How It Works · Data refreshed daily, snapshot 2026-10-07.