SimpleBench: leaderboard

Evaluates 'linguistic adversarial robustness' - ability to answer trick questions. 200+ questions where everyday human reasoning (83.7%) still outperforms frontier LLMs on spatio-temporal reasoning and social intelligence.

Metric: Score (AVG@5). Source: simple-bench.com. Status: saturation imminent. 99 models tracked.

Top models

#ModelScore
1Claude Fable 5.186.6
2Median Human83.7
3Gemini 3.8 Flash82.4
4Claude Fable 581.9
5Muse Spark 1.381.8
6Claude Opus 580.6
7Gemini 3.1 Pro (Preview)79.6
8GPT-5.5 Pro76.9
9Gemini 3.5 Flash76.7
10Gemini 3 Pro (Preview)76.4
11Grok 4.675.9
12Muse Spark 1.274.5
13Qwen 3.7 Max70.4
14Grok 4.570
15GPT-5.569

Interactive version: theaggregate.ai/benchmark?slug=simplebench · How It Works · Data refreshed daily, snapshot 2026-09-05.