StableToolBench — leaderboard

StableToolBench evaluates LLM tool-use systems on solvable tool-query tasks, reporting pass-rate and win-rate scores across instruction, category, and tool subsets.

Metric: Average Pass Rate (%). Source: huggingface.co. 12 models tracked.

Top models

#ModelScore
1Method Name 187.5
2Method Name 277.2
3GPT-4-Turbo-Preview (DFS)73.2
4GPT-3.5-Turbo-1106 (DFS)69.9
5GPT-4-0613 (DFS)69.7
6GPT-3.5-Turbo-0613 (DFS)68.1
7GPT-4-Turbo-Preview (CoT)60.8
8ToolLLaMA v2 (DFS)58.7
9GPT-4-0613 (CoT)55.4
10GPT-3.5-Turbo-1106 (CoT)52.1

Interactive version: theaggregate.ai/benchmark?slug=stabletoolbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.