ToolRobustBench - Runtime Environment: leaderboard

Metric: Success rate (0-1, native; perturbed runtime status and feedback during execution; over ToolRobustBench single-family perturbation instances (clean, light, medium, heavy); binary per-instance tool-call success). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro0.75
2Claude Sonnet 4.60.75
3Qwen Plus0.71
4GPT-5.4 Mini0.71
5Gemini 2.5 Flash0.65
6Gemini 2.5 Pro0.64
7DeepSeek V4 Flash0.62

Interactive version: theaggregate.ai/benchmark?slug=toolrobustbench-runtime-environment · How It Works · Data refreshed daily, snapshot 2026-09-26.