RankJudge - Machine Learning: leaderboard
Metric: Bradley-Terry Elo rating (open scale, mean strength mapped to 1500) of an LLM judge, fitted on joint correctness (the judge must pick the flawed conversation, its flawed turn and its failure type) over RankJudge pairs of multi-turn, document-grounded conversations in which exactly one conversation carries one injected assistant failure; pairs every judge answers alike and the top 5 percent hardest pairs are dropped; machine-learning domain (240 pairs grounded in CS papers); higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 1992 |
| 2 | Claude Sonnet 4.6 | 1755 |
| 3 | Kimi K2.6 | 1743 |
| 4 | Gemma 4 31B | 1708 |
| 5 | Gemini 3 Flash | 1665 |
| 6 | Qwen 3.6 Plus | 1665 |
| 7 | GLM-5.1 | 1665 |
| 8 | Qwen 3.5 397B A17B | 1590 |
| 9 | GPT-5.4 | 1581 |
| 10 | Claude Opus 4.7 | 1523 |
| 11 | Claude Haiku 4.5 | 1492 |
| 12 | Gemma 4 26B A4B | 1428 |
| 13 | Qwen 3.5 122B A10B | 1381 |
| 14 | Qwen 3.5 35B A3B | 1349 |
| 15 | GPT-5.4 Mini | 1343 |
Interactive version: theaggregate.ai/benchmark?slug=rankjudge-machine-learning · How It Works · Data refreshed daily, snapshot 2026-10-07.