BoolQ — leaderboard

Boolean Questions: 15,942 naturally occurring yes/no questions from Google search queries paired with Wikipedia passages. Tests reading comprehension and inference.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 77 models tracked.

Top models

#ModelScore
1Median Human90
2StableBeluga289.4
3falcon-180B89
4GPT-4o Mini (2024-07-18)88.7
5Llama 2 70B Base88.6
6text-davinci-00388.1
7text-davinci-00287.7
8Mistral-7B-v0.187.4
9LLaMA-65B87.1
10GPT-3.5 Turbo (0613)87
11Qwen-14B86.2
12Gemini 1.5 Flash (001)85.8
13Gemma 2 9B85.7
14mpt-30B Instruct85
15Phi-3.5-MoE-instruct84.6

Interactive version: theaggregate.ai/benchmark?slug=boolq · How the rankings work · Data refreshed daily, snapshot 2026-07-22.