Galileo Tool Tasks - ToolACE Single Function Call — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 19 models tracked.

Top models

#ModelScore
1Gemini 1.5 Flash99
2Claude 3.7 Sonnet (20250219)97.5
3O3 Mini (2025-01-31)97.5
4Gemini 2.0 Flash (001)96.5
5GPT-4o (2024-11-20)96.5
6Claude 3.5 Sonnet (20241022)95.5
7Qwen 2.5 72B Instruct95
8O1 (2024-12-17)95
9Gemini 1.5 Pro92.5
10Claude 3.5 Haiku (20241022)90.5
11Llama 3.3 70B Instruct86.5
12GPT-4o Mini83.5
13Mistral Small 377.5
14Mistral Large 2 (Nov) Instruct (2411)72.5
15Llama 3.1 8B Instruct57.5

Interactive version: theaggregate.ai/benchmark?slug=galileo-tool-tasks-toolace-single-function-call · How the rankings work · Data refreshed daily, snapshot 2026-07-22.