LIT-RAGBench (English) - Reasoning: leaderboard

Metric: Accuracy (%) on the 23 Reasoning questions (multi-hop and numerical reasoning over the documents) among the English translation of the 54 questions (translated with GPT-5), judged correct or incorrect against the reference answer by GPT-4.1 (2025-04-14); the generator receives the question with relevant and irrelevant documents of about 512 tokens each, temperature 0 where supported and the maximum reasoning length for reasoning models; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScoreOverall rank
1O387#121
2Qwen 3 235B A22B 2507 (Thinking)87#253 (Qwen 3 235B A22B 2507)
3Gemini 2.5 Flash82.6#237
4GPT-4.1 Mini82.6#346
5Qwen 3 235B A22B 2507 Instruct78.3#291
6O4 Mini78.3#172
7GPT-573.9#91
8GPT-4.173.9#240
9GPT-5 Nano73.9#415
10Gemini 2.5 Pro69.6#145
11GPT-5 Mini69.6#176
12Gemma 3 27B (IT)60.9#509
13Llama 3.3 70B Instruct60.9#520
14Claude Sonnet 456.5#194
15Llama 3.1 8B Instruct26.1#1018

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=lit-ragbench-english-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.