ToolSandbox — leaderboard

ToolSandbox: Evaluates tool calling, API use, function selection, structured arguments, and multi-step tool workflows.

Metric: Avg Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 13 models tracked.

Top models

#ModelScore
1GPT-4o73
2Claude 3 Opus69.2
3GPT-3.5 Turbo65.6
4GPT-464.3
5Claude 3 Sonnet (20240229)63.8
6Gemini 1.5 Pro (001)60.4
7Claude 3 Haiku54.9
8Gemini 1.0 Pro38.1
9Hermes-2-Pro-Mistral-7B31.4
10Mistral 7B Instruct (v0.3)29.8
11c4ai-command-r-v0126.2
12c4ai-command-r-plus24.7

Interactive version: theaggregate.ai/benchmark?slug=toolsandbox · How the rankings work · Data refreshed daily, snapshot 2026-07-22.