SupChain-Bench (Tool Calling, With SOP): leaderboard

Metric: Information retrieval accuracy (%) on SupChain-Bench's 98 order-diagnosis questions answered by calling tools over a simulated supply-chain order database, a question counting as correct when the order information its tool calls retrieve matches what an oracle run of the expert procedure gathers; the prompt adds the expert standard operating procedure, step by step; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini86.73#176
2GPT-584.69#91
3GPT-5 Nano84.69#415
4Claude 3.7 Sonnet40.09#241
5Claude Sonnet 439.79#194
6Gemini 2.5 Pro37.56#145
7Claude Opus 434.69#155
8Claude 3.5 Sonnet29.59#337
9GPT-4.1 Mini28.57#346
10Qwen 3 Max28.57#201
11Gemini 2.5 Flash23.46#237
12DeepSeek R120.2#245
13GPT-4o16.32#333
14GPT-4.116.32#240
15QwQ-32B16.32#410

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=supchain-bench-tool-calling-with-sop · How It Works · Data refreshed daily, snapshot 2026-10-11.