ProLLM - OpenBook Q&A — leaderboard

ProLLM benchmark measuring LLM accuracy on open-book question answering with access to reference documents.

Metric: Score (%). Source: www.prollm.ai. Status: saturation imminent. 85 models tracked.

Top models

#ModelScore
1Grok 494.4
2QwQ-32B91.9
3Grok 3 Mini86.7
4Kimi K286.3
5Command A86.3
6Grok 385.9
7GPT-4.1 Mini85.1
8DeepSeek V385.1
9Gemini 2.5 Pro84.7
10DeepSeek R184.7
11GPT-4.182.7
12O1 Preview81
13Gemini 2.5 Flash80.6
14O180.6
15Qwen 2.5 72B Instruct80.2

Interactive version: theaggregate.ai/benchmark?slug=prollm-openbook-q-a · How the rankings work · Data refreshed daily, snapshot 2026-07-22.