AdaPlanBench - Valid Plan Rate: leaderboard
Metric: Valid plan rate (%): share of tasks that end with a plan satisfying every constraint rather than an early stop or the turn limit, without the rubric gate, 307 household planning tasks with progressively revealed world and user constraints at the mid-level constraint profile: a hidden constraint is disclosed only when a proposed plan violates it (GPT-5.4 constraint judges), and the agent re-plans within the turn budget; temperature 0 (GPT-5 family at its fixed default, mean of three runs); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 91.21 |
| 2 | Gemini 3 Flash | 90.23 |
| 3 | GPT-5 | 89.58 |
| 4 | GPT-5 Mini | 85.34 |
| 5 | Llama 3.3 70B Instruct | 83.71 |
| 6 | Qwen 3 8B | 82.35 |
| 7 | Qwen 3 32B | 80.13 |
| 8 | DeepSeek V4 Flash | 76.97 |
| 9 | Qwen 3 14B | 73.62 |
| 10 | GPT-5 Nano | 67.75 |
Interactive version: theaggregate.ai/benchmark?slug=adaplanbench-valid-plan-rate · How It Works · Data refreshed daily, snapshot 2026-09-29.