EnterpriseBench Corecraft (Held-Out Training Split): leaderboard

Metric: Task pass rate (%) on the 150 held-out Corecraft tasks of the authors' random 1,000/150 training split (the task set available at training time, before tasks were added); a task passes only when an LLM judge finds every rubric criterion satisfied; higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.1 (High)36.86#131 (GPT-5.1)
2Claude Opus 4.533.49#79
3GPT-530.45#91
4Claude Sonnet 4.526.44#138
5GLM-4.625.37#246
6Gemini 3 Pro (Preview)24.84#64
7GPT-5.224.77#105
8Claude Haiku 4.57.21#271

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=enterprisebench-corecraft-held-out-training-split · How It Works · Data refreshed daily, snapshot 2026-10-11.