WildToolBench: leaderboard

Metric: Session accuracy (%): share of the 256 dialogues in which all four tasks are completed correctly, on WildToolBench (256 multi-turn scenarios, four user tasks each, 1,024 tasks, built from real user-log patterns with expert-annotated tool calls), each model through its native function-call format with default decoding; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 57 models tracked.

Top models

#ModelScore
1Gemini 2.0 Flash (Thinking)14.45
2Gemini 2.5 Pro14.06
3Claude Sonnet 412.5
4O112.11
5GLM-4.512.11
6GPT-4o11.72
7Claude 3.7 Sonnet11.33
8Kimi K210.55
9Grok 410.16
10O310.16
11Qwen 3 30B A3B (Non-reasoning)9.77
12DeepSeek R19.38
13DeepSeek V39.38
14Qwen 3 14B (Reasoning)9.38
15Claude Opus 4.19.38

Interactive version: theaggregate.ai/benchmark?slug=wildtoolbench · How It Works · Data refreshed daily, snapshot 2026-10-07.