ExplorationBench - AlienCode: leaderboard

Metric: Held-out accuracy after exploration (%; Best@3: the highest round-4 score among three independent autonomous exploration trajectories of 12 tool calls per round; AlienCode, a calculation language with 31 hidden altered semantics, 70 program-synthesis and prediction tasks graded by an interpreter on private inputs; 70 held-out tasks each answered three times in a tool-disabled copy of the conversation and averaged; highest reasoning setting each API offers). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Claude Opus 5 (Max)87.6
2GPT-5.6 Sol (Max)87.1
3Kimi K3 (Max)77.6
4Gemini 3.8 Flash (High)77.6
5Grok 4.6 (xHigh)67.6
6DeepSeek V4.1 Flash (Max)58.1
7DeepSeek V4 Pro (Max)12.9

Interactive version: theaggregate.ai/benchmark?slug=explorationbench-aliencode · How It Works · Data refreshed daily, snapshot 2026-09-26.