SQuAD 2.0: leaderboard
SQuAD 2.0: Evaluates broad language-model knowledge, reasoning, commonsense, instruction following, or exam-style accuracy.
Metric: F1 (%). Source: rajpurkar.github.io. 179 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Median Human | 74 |
Interactive version: theaggregate.ai/benchmark?slug=squad-2-0 · How It Works · Data refreshed daily, snapshot 2026-09-05.