EnterpriseBench Corecraft (Paper Task Set): leaderboard

Metric: Task pass rate (%) on the Corecraft benchmark task set evaluated in the paper (expert-written customer-support tasks in a simulated PC-parts retailer with 23 MCP tools); a task passes only when an LLM judge finds every expert-authored rubric criterion satisfied across its trajectories; higher is better. Source: arxiv.org. Saturation forecast: Around October 2027. 17 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.6 (Adaptive Reasoning, Max Effort)30.8#60 (Claude Opus 4.6)
2GPT-5.2 (High)29.7#105 (GPT-5.2)
3Gemini 3.1 Pro (Preview)27.2#54
4Claude Opus 4.6 (High)26.2#60 (Claude Opus 4.6)
5DeepSeek V3.2 (Thinking)24.1#198 (DeepSeek V3.2)
6Claude Opus 4.622.1#60
7Grok 4.1 Fast (Reasoning)20.5#208 (Grok 4.1 Fast)
8GPT-5.2 Codex (xHigh)20.1#89 (GPT-5.2 Codex)
9Gemini 3 Flash20#93
10GLM-517.4#137
11Claude Sonnet 4.616.4#85
12Gemini 3 Pro14.4#77
13Qwen 3.5 Plus11.3#123
14Kimi K2.58.7#139
15Mistral Large 33.6#388

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=enterprisebench-corecraft-paper-task-set · How It Works · Data refreshed daily, snapshot 2026-10-11.