ToolFailBench: leaderboard

Metric: Clean tool-use rate (%; share of the 750 tool-required single-turn tasks whose trace calls the needed tool and answers with the returned value, without skipping the tool, ignoring its result or fabricating output; labels by majority vote of a rule classifier and two LLM judges; mock tool returns contradict plausible prior values; temperature 0, 1,024-token output limit; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 19 models tracked.

Top models

#ModelScore
1Grok 4.386.33
2Grok 4.1 Fast (Reasoning)84.11
3Qwen 2.5 32B Instruct82.68
4Qwen 3.6 27B79.33
5Claude Sonnet 4.579.28
6GPT-5.4 Mini79.14
7QwQ-32B79.04
8Qwen 2.5 72B Instruct79
9Qwen 3.6 35B A3B78.47
10Gemma 4 31B (IT)78.12
11Qwen 3.5 27B77.38
12Claude Haiku 4.576.47
13DeepSeek V4 Flash75.84
14GLM-4.7 Flash71.49
15Qwen 3.5 9B70.03

Interactive version: theaggregate.ai/benchmark?slug=toolfailbench · How It Works · Data refreshed daily, snapshot 2026-09-29.