EnterpriseBench Corecraft (Held-Out Training Split): leaderboard
Metric: Task pass rate (%) on the 150 held-out Corecraft tasks of the authors' random 1,000/150 training split (the task set available at training time, before tasks were added); a task passes only when an LLM judge finds every rubric criterion satisfied; higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.1 (High) | 36.86 | #131 (GPT-5.1) |
| 2 | Claude Opus 4.5 | 33.49 | #79 |
| 3 | GPT-5 | 30.45 | #91 |
| 4 | Claude Sonnet 4.5 | 26.44 | #138 |
| 5 | GLM-4.6 | 25.37 | #246 |
| 6 | Gemini 3 Pro (Preview) | 24.84 | #64 |
| 7 | GPT-5.2 | 24.77 | #105 |
| 8 | Claude Haiku 4.5 | 7.21 | #271 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=enterprisebench-corecraft-held-out-training-split · How It Works · Data refreshed daily, snapshot 2026-10-11.