WildToolBench - Multi-Tool Tasks: leaderboard

Metric: Accuracy (%) on compositional tasks needing multi-step sequential or parallel tool calls, on WildToolBench (256 multi-turn scenarios, four user tasks each, 1,024 tasks, built from real user-log patterns with expert-annotated tool calls), each model through its native function-call format with default decoding; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 57 models tracked.

Top models

#ModelScore
1GPT-4.144.14
2Claude Sonnet 443.75
3GPT-4o41.8
4Grok 441.41
5DeepSeek R141.02
6DeepSeek V3.140.63
7GLM-4.540.63
8Seed-1.640.23
9Gemini 2.0 Flash (Thinking)40.23
10Claude Opus 4.139.84
11O339.45
12Claude 3.7 Sonnet39.06
13O139.06
14DeepSeek V338.67
15Qwen 2.5 32B Instruct38.67

Interactive version: theaggregate.ai/benchmark?slug=wildtoolbench-multi-tool-tasks · How It Works · Data refreshed daily, snapshot 2026-10-07.