SupChain-Bench (Tool Calling): leaderboard

Metric: Information retrieval accuracy (%) on SupChain-Bench's 98 order-diagnosis questions answered by calling tools over a simulated supply-chain order database, a question counting as correct when the order information its tool calls retrieve matches what an oracle run of the expert procedure gathers; the model sees only the question and the tool schemas; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini46.93#176
2GPT-535.71#91
3Claude 3.7 Sonnet35.71#241
4Claude Sonnet 431.63#194
5Claude Opus 426.53#155
6GPT-5 Nano25.51#415
7GPT-4o20.4#333
8DeepSeek R117.14#245
9GPT-4.112.24#240
10Kimi K212.24#236
11Gemini 2.5 Pro11.22#145
12GPT-4.1 Mini11.22#346
13Claude 3.5 Sonnet11.22#337
14Gemini 2.5 Flash10.2#237
15DeepSeek V310.2#312

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=supchain-bench-tool-calling · How It Works · Data refreshed daily, snapshot 2026-10-11.