AdaPlanBench - Valid Plan Rate: leaderboard

Metric: Valid plan rate (%): share of tasks that end with a plan satisfying every constraint rather than an early stop or the turn limit, without the rubric gate, 307 household planning tasks with progressively revealed world and user constraints at the mid-level constraint profile: a hidden constraint is disclosed only when a proposed plan violates it (GPT-5.4 constraint judges), and the agent re-plans within the turn budget; temperature 0 (GPT-5 family at its fixed default, mean of three runs); higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)91.21
2Gemini 3 Flash90.23
3GPT-589.58
4GPT-5 Mini85.34
5Llama 3.3 70B Instruct83.71
6Qwen 3 8B82.35
7Qwen 3 32B80.13
8DeepSeek V4 Flash76.97
9Qwen 3 14B73.62
10GPT-5 Nano67.75

Interactive version: theaggregate.ai/benchmark?slug=adaplanbench-valid-plan-rate · How It Works · Data refreshed daily, snapshot 2026-09-29.