IRB1K (Closed-Book) - Cross-Lingual: leaderboard

Metric: Correctness (%, 0-100): share of answers judged correct (mean of two judges, GPT-4.1-mini and Qwen3-Next-80B, each labelling correct, incorrect or not attempted) on the 252 valid-premise questions whose supporting evidence is in another language (English question, cross-lingual evidence), of IRB1K's 1,000 short-answer factuality questions generated from citing sentences of 2024-2025 Wikipedia articles; closed-book: no retrieved context; the prompt sets the current date to 29 September 2025, invites 'I don't know' and false-premise answers, and runs once; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 (Medium)31.5#91 (GPT-5)
2GPT-4.121.4#240
3GPT-OSS-120B11.9#330
4GPT-4.1 Mini9.7#346
5Llama 3.3 70B Instruct9.5#520
6GPT-5 Mini (Medium)9.1#176 (GPT-5 Mini)
7DeepSeek R18.5#245
8Llama 4 Scout4.8#646

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=irb1k-closed-book-cross-lingual · How It Works · Data refreshed daily, snapshot 2026-10-11.