EnigmaForge - Intuition: leaderboard

Metric: Task success (%) on implicit-condition instances, where the question is not stated and must be discovered from the story. Source: robottwo.github.io. Saturation forecast: Around December 2026. 22 models tracked.

Top models

#ModelScore
1GPT-6 (Low)80.1
2GPT-6 Sol (Low)75
3GPT-5.6 Sol (Low)54.2
4Gemini 3.7 Flash (Low)52.9
5Gemini 3.8 Flash (Low)39.3
6GLM-5.3 (Low)37.7
7GPT-5.6 Terra (Low)24.2
8GPT-6 Luna (Low)22.9
9Claude Opus 5.5 (Low)22
10Grok 4.6 (Low)21.2
11GPT-5.6 Luna (Low)16.7
12Qwen 3.8 27B (Low)12.1

Interactive version: theaggregate.ai/benchmark?slug=enigmaforge-intuition · How It Works · Data refreshed daily, snapshot 2026-09-29.