ToolBench - Google Sheets — leaderboard

Metric: Task Score. Source: huggingface.co. 44 models tracked.

Top models

#ModelScore
1CodeLlama-34B-hf64.29
2GPT-462.9
3text-davinci-00362.9
4CodeLlama-34B-Instruct-hf61.11
5Llama 2 70B58.57
6CodeLlama-34B-Python-hf55.87
7GPT-3.5 Turbo51.4
8CodeLlama-13B-hf51.26
9CodeLlama-13B-Python-hf50.79
10CodeLlama-7B-Python-hf49.13
11CodeLlama-13B-Instruct-hf48.97
12starcoder48
13CodeLlama-7B-Instruct-hf41.27
14CodeLlama-7B-hf38.08
15LLaMA-30B37.1

Interactive version: theaggregate.ai/benchmark?slug=toolbench-google-sheets · How the rankings work · Data refreshed daily, snapshot 2026-07-22.