ToolBench Leaderboard — leaderboard

ToolBench action-generation leaderboard measuring how well models manipulate APIs and tools across weather, search, booking, spreadsheet, WebShop, and tabletop tasks.

Metric: Average Task Score. Source: huggingface.co. Status: saturation imminent. 44 models tracked.

Top models

#ModelScore
1GPT-468.76
2text-davinci-00367.24
3CodeLlama-34B-Instruct-hf64.79
4CodeLlama-34B-hf62.94
5Llama 2 70B61.03
6CodeLlama-13B-Instruct-hf60.49
7CodeLlama-34B-Python-hf59.16
8CodeLlama-13B-hf56.89
9GPT-3.5 Turbo56.65
10CodeLlama-13B-Python-hf56.31
11LLaMA-65B55.59
12CodeLlama-7B-Python-hf52.15
13CodeLlama-7B-Instruct-hf50.48
14starcoder49.75
15LLaMA-30B49.59

Interactive version: theaggregate.ai/benchmark?slug=toolbench-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.