ToolBench - Tabletop — leaderboard

Metric: Task Score. Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1GPT-481
2text-davinci-00366.7
3CodeLlama-13B-Instruct-hf57.62
4CodeLlama-34B-hf51.32
5CodeLlama-34B-Instruct-hf47.46
6Llama 2 70B45.4
7CodeLlama-13B-Python-hf41.53
8CodeLlama-13B-hf40.95
9LLaMA-30B34.3
10CodeLlama-34B-Python-hf33.33
11GPT-3.5 Turbo33.3
12LLaMA-65B30.5
13Llama 2 13B23.81
14CodeLlama-7B-Python-hf22.86
15starcoder21.9

Interactive version: theaggregate.ai/benchmark?slug=toolbench-tabletop · How the rankings work · Data refreshed daily, snapshot 2026-07-22.