EnigmaEval: leaderboard

EnigmaEval is a benchmark from puzzle hunts, testing AI with complex reasoning, creative problem-solving, and cross-domain knowledge synthesis.

Metric: Score. Source: scale.com. Status: saturation imminent. 5 models tracked.

Top models

#ModelScore
1Claude Fable 5 (High)39.28
2GPT-5.6 Sol (High)37.12
3Gemini 3.1 Pro (Preview) (High)36.78
4Gemini 3.5 Flash (High)25.41
5Claude Opus 4.8 (xHigh)23.51

Interactive version: theaggregate.ai/benchmark?slug=enigmaeval · How It Works · Data refreshed daily, snapshot 2026-09-05.