WildToolBench - Single-Tool Tasks: leaderboard

Metric: Accuracy (%) on tasks solvable with one tool call, on WildToolBench (256 multi-turn scenarios, four user tasks each, 1,024 tasks, built from real user-log patterns with expert-annotated tool calls), each model through its native function-call format with default decoding; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 57 models tracked.

Top models

#ModelScore
1O361.72
2GPT-4o60.16
3Claude Sonnet 460.16
4Qwen 3 8B (Non-reasoning)60.16
5Qwen 3 4B (Reasoning)60.16
6Doubao-1.5-Thinking-Pro60.16
7Grok 459.38
8Qwen 2.5 72B Instruct58.98
9DeepSeek V358.98
10Doubao-1.5-Pro58.59
11Claude 3.7 Sonnet57.81
12Qwen 3 32B (Non-reasoning)57.81
13GLM-4.557.81
14GPT-4.157.42
15Doubao-Seed-1.6 (Thinking)57.42

Interactive version: theaggregate.ai/benchmark?slug=wildtoolbench-single-tool-tasks · How It Works · Data refreshed daily, snapshot 2026-10-07.