RankJudge - Finance: leaderboard

Metric: Bradley-Terry Elo rating (open scale, mean strength mapped to 1500) of an LLM judge, fitted on joint correctness (the judge must pick the flawed conversation, its flawed turn and its failure type) over RankJudge pairs of multi-turn, document-grounded conversations in which exactly one conversation carries one injected assistant failure; pairs every judge answers alike and the top 5 percent hardest pairs are dropped; finance domain (218 pairs grounded in S&P 500 10-K filings); higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)1933
2Gemini 3 Flash1713
3Gemma 4 31B1662
4GPT-5.41653
5Kimi K2.61634
6Claude Sonnet 4.61624
7Qwen 3.6 Plus1597
8GLM-5.11546
9Claude Haiku 4.51490
10Qwen 3.5 397B A17B1490
11Claude Opus 4.71395
12Qwen 3.5 122B A10B1395
13Gemma 4 26B A4B1380
14GPT-5.4 Mini1279
15Qwen 3.5 35B A3B1272

Interactive version: theaggregate.ai/benchmark?slug=rankjudge-finance · How It Works · Data refreshed daily, snapshot 2026-10-07.