IRB1K (Closed-Book) - English: leaderboard

Metric: Correctness (%, 0-100): share of answers judged correct (mean of two judges, GPT-4.1-mini and Qwen3-Next-80B, each labelling correct, incorrect or not attempted) on the 548 valid-premise questions whose supporting evidence is in English, of IRB1K's 1,000 short-answer factuality questions generated from citing sentences of 2024-2025 Wikipedia articles; closed-book: no retrieved context; the prompt sets the current date to 29 September 2025, invites 'I don't know' and false-premise answers, and runs once; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 (Medium)41.9#91 (GPT-5)
2GPT-4.130.6#240
3GPT-OSS-120B16.3#330
4Llama 3.3 70B Instruct14.8#520
5GPT-5 Mini (Medium)14.8#176 (GPT-5 Mini)
6GPT-4.1 Mini11.4#346
7DeepSeek R111.3#245
8Llama 4 Scout8.5#646

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=irb1k-closed-book-english · How It Works · Data refreshed daily, snapshot 2026-10-11.