IRB1K (RAG) - English: leaderboard

Metric: Correctness (%, 0-100): share of answers judged correct (mean of two judges, GPT-4.1-mini and Qwen3-Next-80B, each labelling correct, incorrect or not attempted) on the 548 valid-premise questions whose supporting evidence is in English, of IRB1K's 1,000 short-answer factuality questions generated from citing sentences of 2024-2025 Wikipedia articles; RAG: the top-5 documents retrieved by text-embedding-3-small from the IRB1K corpus (512-token chunks, documents ranked by their best chunk) are given with their source and publication date; the prompt sets the current date to 29 September 2025, invites 'I don't know' and false-premise answers, and runs once; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4.184.9#240
2GPT-5 (Medium)84.9#91 (GPT-5)
3GPT-OSS-120B84.4#330
4GPT-5 Mini (Medium)83.1#176 (GPT-5 Mini)
5GPT-4.1 Mini82.8#346
6DeepSeek R179.9#245
7Llama 3.3 70B Instruct78.6#520
8Llama 4 Scout76#646

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=irb1k-rag-english · How It Works · Data refreshed daily, snapshot 2026-10-11.