Scale AI - EnigmaEval: leaderboard

Complex, multi-step reasoning tasks requiring models to chain together multiple inference steps to reach correct solutions.

Metric: Score. Source: scale.com. Status: years away from saturation. 5 models tracked.

Top models

#ModelScore
1Claude Fable 5 (High)39.28
2GPT-5.6 Sol (High)37.12
3Gemini 3.1 Pro (Preview) (High)36.78
4Gemini 3.5 Flash (High)25.41
5Claude Opus 4.8 (xHigh)23.51

Interactive version: theaggregate.ai/benchmark?slug=scale-ai-enigmaeval · How It Works · Data refreshed daily, snapshot 2026-09-05.