ToolRobustBench - Tool Output: leaderboard

Metric: Success rate (0-1, native; noisy, missing or conflicting returned evidence after the call; over ToolRobustBench single-family perturbation instances (clean, light, medium, heavy); binary per-instance tool-call success). Source: arxiv.org. Saturation forecast: Around March 2027. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.60.57
2DeepSeek V4 Pro0.51
3GPT-5.4 Mini0.5
4Gemini 2.5 Pro0.48
5Gemini 2.5 Flash0.42
6Qwen Plus0.38
7DeepSeek V4 Flash0.32

Interactive version: theaggregate.ai/benchmark?slug=toolrobustbench-tool-output · How It Works · Data refreshed daily, snapshot 2026-09-26.