EnigmaEval — leaderboard

EnigmaEval is a benchmark from puzzle hunts, testing AI with complex reasoning, creative problem-solving, and cross-domain knowledge synthesis.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 39 models tracked.

Top models

#ModelScore
1GPT-5.4 Pro (xHigh)23.82
2Gemini 3.1 Pro (Preview)19.76
3GPT-5 Pro18.75
4GPT-5.4 (xHigh)15.96
5O3 (Medium)13.09
6Claude Opus 4.511.91
7GPT-5.111.23
8GPT-510.47
9GPT-5.210.39
10O4 Mini (High)9.21
11GPT-5 Mini8.19
12Claude Opus 4.6 (Max)7.6
13Claude Opus 4.17.18
14O1 Pro6.14
15Claude Sonnet 4.56

Interactive version: theaggregate.ai/benchmark?slug=enigmaeval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.