StableToolBench — leaderboard
StableToolBench evaluates LLM tool-use systems on solvable tool-query tasks, reporting pass-rate and win-rate scores across instruction, category, and tool subsets.
Metric: Average Pass Rate (%). Source: huggingface.co. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Method Name 1 | 87.5 |
| 2 | Method Name 2 | 77.2 |
| 3 | GPT-4-Turbo-Preview (DFS) | 73.2 |
| 4 | GPT-3.5-Turbo-1106 (DFS) | 69.9 |
| 5 | GPT-4-0613 (DFS) | 69.7 |
| 6 | GPT-3.5-Turbo-0613 (DFS) | 68.1 |
| 7 | GPT-4-Turbo-Preview (CoT) | 60.8 |
| 8 | ToolLLaMA v2 (DFS) | 58.7 |
| 9 | GPT-4-0613 (CoT) | 55.4 |
| 10 | GPT-3.5-Turbo-1106 (CoT) | 52.1 |
Interactive version: theaggregate.ai/benchmark?slug=stabletoolbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.