Agent Planning Benchmark - Tool-Broken Recovery: leaderboard

Metric: Replace rate (%): share of 300 cases in which, after a critical tool returns an error, the model switches to the provided substitute tool (the optimal behaviour, versus altering, retrying or refusing); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview) (High)78
2Qwen 3 VL 235B A22B Instruct77.7
3Claude Sonnet 4.576.7
4GPT-576.3
5Gemini 2.5 Pro74.7
6Qwen 3 VL 32B Instruct68.7
7Qwen 3 VL 30B A3B Instruct55.7
8Gemini 2.5 Flash (Non-reasoning)44.3
9GPT-4o42.3

Interactive version: theaggregate.ai/benchmark?slug=agent-planning-benchmark-tool-broken-recovery · How It Works · Data refreshed daily, snapshot 2026-09-29.