UserToolBench - Multi-Tool: leaderboard

Metric: Exact trajectory accuracy (%; multi-tool tasks; a trajectory counts when it matches the profile-conditioned reference: the tool name and decision-relevant arguments, the ordered call sequence for multi-tool tasks, or the correct clarify-or-infer decision when information is missing; 799 profile-hidden evaluation instances from long-horizon interaction trajectories of 10 privacy-sanitized user profiles over 36 tool sets; the model sees the interaction history, the current request and the tool schemas but not the profile). Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1GLM-538.41
2Qwen 3.6 Plus33.51
3DeepSeek V4 Pro33.26
4Gemini 3.5 Flash30.73
5Kimi K2.626.44
6GPT-5.413.48

Interactive version: theaggregate.ai/benchmark?slug=usertoolbench-multi-tool · How It Works · Data refreshed daily, snapshot 2026-09-29.