Agent Planning Benchmark - Step-wise Planning: leaderboard

Metric: Correctness rate (%) of the next 1 to 3 actions predicted from a partially executed trajectory with observed tool returns (900 instances); binary plan correctness judged by Gemini 3 Pro against a reference trajectory (92.2% agreement with human annotation); multimodal tool-use cases aggregated from OpenCUA, GTA, GAIA, ToolBench, FrameThinker and real agent-platform traffic; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview) (High)80.1
2GPT-577.4
3Claude Sonnet 4.577.3
4Gemini 2.5 Pro69.2
5Gemini 2.5 Flash (Non-reasoning)66.6
6Qwen 3 VL 235B A22B Instruct60.2
7Qwen 3 VL 32B Instruct59.9
8GPT-4o57.6
9Qwen 3 VL 30B A3B Instruct49.6

Interactive version: theaggregate.ai/benchmark?slug=agent-planning-benchmark-step-wise-planning · How It Works · Data refreshed daily, snapshot 2026-09-29.