ToolRobustBench (Overall): leaderboard

Metric: Success rate (0-1, native; mean tool-call success across the four perturbation families and three severities; over ToolRobustBench single-family perturbation instances (clean, light, medium, heavy); binary per-instance tool-call success). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.60.77
2DeepSeek V4 Pro0.75
3GPT-5.4 Mini0.71
4Qwen Plus0.71
5Gemini 2.5 Pro0.69
6DeepSeek V4 Flash0.67
7Gemini 2.5 Flash0.66

Interactive version: theaggregate.ai/benchmark?slug=toolrobustbench-overall · How It Works · Data refreshed daily, snapshot 2026-09-26.