PlanningBench - Avg-pass: leaderboard
Metric: Average checklist pass rate (%): share of an instance's verification checklist items (general, task-specific and stateful constraints) the plan satisfies, averaged over the PlanningBench evaluation instances; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 (xHigh) | 92.35 |
| 2 | GPT-5.4 (Medium) | 90.03 |
| 3 | Gemini 3.1 Pro (Preview) | 88.36 |
| 4 | GPT-5.4 (High) | 84.6 |
| 5 | Seed 2.0 Pro (High) | 84.02 |
| 6 | DeepSeek V3.2 (Thinking) | 80.21 |
| 7 | Hy3-preview | 78.14 |
| 8 | Qwen 3.5 Plus (Thinking) | 77.1 |
| 9 | Seed 2.0 Pro (Medium) | 77.09 |
| 10 | DeepSeek V3.2 Exp | 76.69 |
| 11 | Gemini 2.5 Pro | 76.41 |
| 12 | DeepSeek R1 | 74.04 |
| 13 | Seed 1.8 | 71.69 |
| 14 | Qwen 3 30B A3B | 57.4 |
| 15 | Qwen 3 32B | 30.11 |
Interactive version: theaggregate.ai/benchmark?slug=planningbench-avg-pass · How It Works · Data refreshed daily, snapshot 2026-10-07.