BoolQ: leaderboard

Boolean Questions: 15,942 naturally occurring yes/no questions from Google search queries paired with Wikipedia passages. Tests reading comprehension and inference.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 77 models tracked.

Top models

#ModelScore
1Median Human90
2GPT-4o Mini (2024-07-18)88.7
3Llama 2 70B Base88.6
4text-davinci-00388.1
5text-davinci-00287.7
6Mistral-7B-v0.187.4
7LLaMA-65B87.1
8GPT-3.5 Turbo (0613)87
9Qwen-14B86.2
10Gemini 1.5 Flash (001)85.8
11Gemma 2 9B85.7
12mpt-30B Instruct85
13Phi-3.5-MoE-instruct84.6
14gemma-7B83.2
15Mistral 7B Instruct (v0.2)83.2

Interactive version: theaggregate.ai/benchmark?slug=boolq · How It Works · Data refreshed daily, snapshot 2026-09-05.