Agent Planning Benchmark - Holistic Planning: leaderboard

Metric: Correctness rate (%) of complete plans and tool chains produced in one pass for 1,109 long-horizon instances; binary plan correctness judged by Gemini 3 Pro against a reference trajectory (92.2% agreement with human annotation); multimodal tool-use cases aggregated from OpenCUA, GTA, GAIA, ToolBench, FrameThinker and real agent-platform traffic; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1GPT-574.5
2Gemini 3 Pro (Preview) (High)71.3
3Claude Sonnet 4.564.1
4Gemini 2.5 Pro55
5Gemini 2.5 Flash (Non-reasoning)36.5
6Qwen 3 VL 32B Instruct22.5
7Qwen 3 VL 235B A22B Instruct21.8
8GPT-4o19.5
9Qwen 3 VL 30B A3B Instruct10.7

Interactive version: theaggregate.ai/benchmark?slug=agent-planning-benchmark-holistic-planning · How It Works · Data refreshed daily, snapshot 2026-09-29.