ToolSandbox: leaderboard

ToolSandbox: Evaluates tool calling, API use, function selection, structured arguments, and multi-step tool workflows.

Metric: Avg Score (self-reported). Source: benchmarklist.com. Status: saturated. 13 models tracked.

Top models

#ModelScore
1GPT-4o (2024-05-13)73
2Claude 3 Opus (20240229)69.2
3GPT-3.5 Turbo (0125)65.6
4GPT-4 Preview (0125)64.3
5Claude 3 Sonnet (20240229)63.8
6Gemini 1.5 Pro (001)60.4
7Claude 3 Haiku (20240307)54.9
8Gemini 1.0 Pro38.1
9Hermes-2-Pro-Mistral-7B31.4
10Mistral 7B Instruct (v0.3)29.8
11c4ai-command-r-v0126.2
12c4ai-command-r-plus24.7

Interactive version: theaggregate.ai/benchmark?slug=toolsandbox · How It Works · Data refreshed daily, snapshot 2026-09-05.