OpenEval - BoolQ — leaderboard

Metric: Exact Match (%). Source: huggingface.co. 22 models tracked.

Top models

#ModelScore
1Llama 2 70B87.5
2falcon-40B87.5
3LLaMA-65B84.38
4LLaMA-13B84.38
5Llama 2 7B84.38
6falcon-40B Instruct84.38
7LLaMA-30B82.81
8Llama 2 13B81.25
9falcon-7B Instruct79.69
10LLaMA-7B75
11mpt-30B75
12RedPajama-INCITE-Base-3B-v175
13falcon-7B73.44
14pythia-6.9B70.31
15RedPajama-INCITE-Instruct-3B-v168.75

Interactive version: theaggregate.ai/benchmark?slug=openeval-boolq · How the rankings work · Data refreshed daily, snapshot 2026-07-22.