RankJudge - Biomedicine: leaderboard

Metric: Bradley-Terry Elo rating (open scale, mean strength mapped to 1500) of an LLM judge, fitted on joint correctness (the judge must pick the flawed conversation, its flawed turn and its failure type) over RankJudge pairs of multi-turn, document-grounded conversations in which exactly one conversation carries one injected assistant failure; pairs every judge answers alike and the top 5 percent hardest pairs are dropped; biomedical domain (194 pairs grounded in PubMedQA contexts); higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)1910
2Gemini 3 Flash1759
3Kimi K2.61732
4Qwen 3.6 Plus1719
5Claude Sonnet 4.61707
6GLM-5.11671
7Gemma 4 31B1671
8GPT-5.41617
9Qwen 3.5 397B A17B1597
10Claude Opus 4.71587
11Claude Haiku 4.51487
12Gemma 4 26B A4B1487
13Qwen 3.5 35B A3B1471
14Qwen 3.5 122B A10B1454
15DeepSeek V3.21370

Interactive version: theaggregate.ai/benchmark?slug=rankjudge-biomedicine · How It Works · Data refreshed daily, snapshot 2026-10-07.