ExplorationBench - AlienLogic: leaderboard
Metric: Held-out accuracy after exploration (%; Best@3: the highest round-4 score among three independent autonomous exploration trajectories of 12 proofs per round; AlienLogic, a first-order natural-deduction system with 24 patched inference rules, 70 tasks checked by a proof checker, designated unprovable ones credited only when the system declines to prove them; 70 held-out tasks each answered three times in a tool-disabled copy of the conversation and averaged; highest reasoning setting each API offers). Source: arxiv.org. Saturation forecast: Around May 2027. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4.6 (xHigh) | 83.8 |
| 2 | Claude Opus 5 (Max) | 77.6 |
| 3 | GPT-5.6 Sol (Max) | 75.2 |
| 4 | DeepSeek V4 Pro (Max) | 73.8 |
| 5 | Kimi K3 (Max) | 72.9 |
| 6 | DeepSeek V4.1 Flash (Max) | 72.4 |
| 7 | Gemini 3.8 Flash (High) | 58.1 |
Interactive version: theaggregate.ai/benchmark?slug=explorationbench-alienlogic · How It Works · Data refreshed daily, snapshot 2026-09-26.