RankJudge - Finance: leaderboard
Metric: Bradley-Terry Elo rating (open scale, mean strength mapped to 1500) of an LLM judge, fitted on joint correctness (the judge must pick the flawed conversation, its flawed turn and its failure type) over RankJudge pairs of multi-turn, document-grounded conversations in which exactly one conversation carries one injected assistant failure; pairs every judge answers alike and the top 5 percent hardest pairs are dropped; finance domain (218 pairs grounded in S&P 500 10-K filings); higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 1933 |
| 2 | Gemini 3 Flash | 1713 |
| 3 | Gemma 4 31B | 1662 |
| 4 | GPT-5.4 | 1653 |
| 5 | Kimi K2.6 | 1634 |
| 6 | Claude Sonnet 4.6 | 1624 |
| 7 | Qwen 3.6 Plus | 1597 |
| 8 | GLM-5.1 | 1546 |
| 9 | Claude Haiku 4.5 | 1490 |
| 10 | Qwen 3.5 397B A17B | 1490 |
| 11 | Claude Opus 4.7 | 1395 |
| 12 | Qwen 3.5 122B A10B | 1395 |
| 13 | Gemma 4 26B A4B | 1380 |
| 14 | GPT-5.4 Mini | 1279 |
| 15 | Qwen 3.5 35B A3B | 1272 |
Interactive version: theaggregate.ai/benchmark?slug=rankjudge-finance · How It Works · Data refreshed daily, snapshot 2026-10-07.