ToolRobustBench - Tool Interface: leaderboard

Metric: Success rate (0-1, native; interface-level perturbations to tool names, descriptions and boundaries; over ToolRobustBench single-family perturbation instances (clean, light, medium, heavy); binary per-instance tool-call success). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.60.97
2DeepSeek V4 Flash0.96
3Qwen Plus0.93
4DeepSeek V4 Pro0.93
5Gemini 2.5 Pro0.9
6Gemini 2.5 Flash0.89
7GPT-5.4 Mini0.86

Interactive version: theaggregate.ai/benchmark?slug=toolrobustbench-tool-interface · How It Works · Data refreshed daily, snapshot 2026-09-26.