SimpleBench — leaderboard

Evaluates 'linguistic adversarial robustness' - ability to answer trick questions. 200+ questions where everyday human reasoning (83.7%) still outperforms frontier LLMs on spatio-temporal reasoning and social intelligence.

Metric: Score (AVG@5). Source: simple-bench.com. Status: saturation imminent. 87 models tracked.

Top models

#ModelScore
1Human Expert95.4
2Median Human83.7
3Claude Fable 581.9
4Gemini 3.1 Pro (Preview)79.6
5GPT-5.5 Pro76.9
6Gemini 3.5 Flash76.7
7Gemini 3 Pro (Preview)76.4
8Qwen 3.7 Max70.4
9Grok 4.570
10GPT-5.569
11Claude Opus 4.667.6
12Claude Opus 4.864.8
13GPT-5.6 Sol (xHigh)64.8
14Qwen 3.6 Max Preview63
15Gemini 2.5 Pro (Preview 06-05)62.4

Interactive version: theaggregate.ai/benchmark?slug=simplebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.