RankJudge - Biomedicine: leaderboard
Metric: Bradley-Terry Elo rating (open scale, mean strength mapped to 1500) of an LLM judge, fitted on joint correctness (the judge must pick the flawed conversation, its flawed turn and its failure type) over RankJudge pairs of multi-turn, document-grounded conversations in which exactly one conversation carries one injected assistant failure; pairs every judge answers alike and the top 5 percent hardest pairs are dropped; biomedical domain (194 pairs grounded in PubMedQA contexts); higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 1910 |
| 2 | Gemini 3 Flash | 1759 |
| 3 | Kimi K2.6 | 1732 |
| 4 | Qwen 3.6 Plus | 1719 |
| 5 | Claude Sonnet 4.6 | 1707 |
| 6 | GLM-5.1 | 1671 |
| 7 | Gemma 4 31B | 1671 |
| 8 | GPT-5.4 | 1617 |
| 9 | Qwen 3.5 397B A17B | 1597 |
| 10 | Claude Opus 4.7 | 1587 |
| 11 | Claude Haiku 4.5 | 1487 |
| 12 | Gemma 4 26B A4B | 1487 |
| 13 | Qwen 3.5 35B A3B | 1471 |
| 14 | Qwen 3.5 122B A10B | 1454 |
| 15 | DeepSeek V3.2 | 1370 |
Interactive version: theaggregate.ai/benchmark?slug=rankjudge-biomedicine · How It Works · Data refreshed daily, snapshot 2026-10-07.