SimpleQA Verified: leaderboard

Factual accuracy benchmark testing short factoid questions with verifiable answers. Measures how often models give correct, non-hallucinated responses to straightforward knowledge queries.

Metric: Accuracy (%). Source: epoch.ai. Status: years away from saturation. 78 models tracked.

Top models

#ModelScore
1GPT-6 (Max)75.6
2Gemini 3.1 Pro (Preview) (High)73.5
3Claude Fable 5.1 (Max)70.8
4GPT-5.6 Sol (Max)69.7
5Gemini 3.7 Flash (High)69.2
6Gemini 3 Flash (Preview) (High)66.8
7Gemini 3.5 Flash (High)66.2
8Gemini 3.6 Flash (High)66.2
9GPT-5.5 (xHigh)63
10Muse Spark 1.2 (xHigh)60.3
11Claude Opus 5 (Max)59.9
12Muse Spark 1.157.79
13Qwen 3.7 Max55.77
14Claude Opus 4.8 (Max)53
15DeepSeek V4 Pro (0813) (Max)52.91

Interactive version: theaggregate.ai/benchmark?slug=simpleqa-verified · How It Works · Data refreshed daily, snapshot 2026-09-05.