WildToolBench - Clarification Tasks: leaderboard

Metric: Accuracy (%) on tasks where the correct response is to ask the user for missing information, on WildToolBench (256 multi-turn scenarios, four user tasks each, 1,024 tasks, built from real user-log patterns with expert-annotated tool calls), each model through its native function-call format with default decoding; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 57 models tracked.

Top models

#ModelScore
1Gemini 2.0 Flash (Thinking)52.34
2O148.05
3Gemini 2.5 Pro46.88
4O344.92
5GLM-4.544.53
6DeepSeek R143.75
7Claude Sonnet 441.8
8Claude Opus 4.141.8
9Claude 3.7 Sonnet41.41
10Qwen 3 30B A3B (Non-reasoning)41.41
11Kimi K239.84
12Qwen 3 32B (Non-reasoning)39.84
13Qwen 3 14B (Non-reasoning)39.84
14Qwen 3 8B (Thinking)39.84
15GPT-4o39.45

Interactive version: theaggregate.ai/benchmark?slug=wildtoolbench-clarification-tasks · How It Works · Data refreshed daily, snapshot 2026-10-07.