SimpleQA Verified — leaderboard

Factual accuracy benchmark testing short factoid questions with verifiable answers. Measures how often models give correct, non-hallucinated responses to straightforward knowledge queries.

Metric: Accuracy (%). Source: epoch.ai. Status: saturation imminent. 65 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)77.3
2Gemini 3 Pro (Preview)72.9
3GPT-5.6 Sol (Max)71.6
4Gemini 3.5 Flash (High)68.4
5Qwen 3 Max (2025-09-23)67.47
6Gemini 3 Flash (Preview)67.4
7Muse Spark66.3
8GPT-5.5 Pro (xHigh)64.5
9GPT-5.5 (xHigh)63.1
10Qwen 3.7 Max58.52
11DeepSeek V4 Pro (Max)57
12Qwen 3.6 Max Preview56.93
13Gemini 2.5 Pro56
14Grok 4.5 (High)53.5
15O3 (2025-04-16) (High)53

Interactive version: theaggregate.ai/benchmark?slug=simpleqa-verified · How the rankings work · Data refreshed daily, snapshot 2026-07-22.