ExplorationBench - AlienLogic: leaderboard

Metric: Held-out accuracy after exploration (%; Best@3: the highest round-4 score among three independent autonomous exploration trajectories of 12 proofs per round; AlienLogic, a first-order natural-deduction system with 24 patched inference rules, 70 tasks checked by a proof checker, designated unprovable ones credited only when the system declines to prove them; 70 held-out tasks each answered three times in a tool-disabled copy of the conversation and averaged; highest reasoning setting each API offers). Source: arxiv.org. Saturation forecast: Around May 2027. 10 models tracked.

Top models

#ModelScore
1Grok 4.6 (xHigh)83.8
2Claude Opus 5 (Max)77.6
3GPT-5.6 Sol (Max)75.2
4DeepSeek V4 Pro (Max)73.8
5Kimi K3 (Max)72.9
6DeepSeek V4.1 Flash (Max)72.4
7Gemini 3.8 Flash (High)58.1

Interactive version: theaggregate.ai/benchmark?slug=explorationbench-alienlogic · How It Works · Data refreshed daily, snapshot 2026-09-26.