Agent Planning Benchmark - Tool-Extraneous Planning: leaderboard

Metric: Correctness rate (%) when 2, 4, 6, 8 or 10 semantically similar but irrelevant tools are added to the tool set, over 1,500 instances (750 holistic and 750 step-wise, 150 per distractor level); binary plan correctness judged by Gemini 3 Pro against a reference trajectory (92.2% agreement with human annotation); multimodal tool-use cases aggregated from OpenCUA, GTA, GAIA, ToolBench, FrameThinker and real agent-platform traffic; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.576.4
2Gemini 3 Pro (Preview) (High)76.4
3GPT-567
4Gemini 2.5 Pro66.8
5Gemini 2.5 Flash (Non-reasoning)59.8
6Qwen 3 VL 235B A22B Instruct52.6
7Qwen 3 VL 32B Instruct51.9
8GPT-4o45.5
9Qwen 3 VL 30B A3B Instruct40.2

Interactive version: theaggregate.ai/benchmark?slug=agent-planning-benchmark-tool-extraneous-planning · How It Works · Data refreshed daily, snapshot 2026-09-29.