LPS-Bench - Inefficient Planning and Resource Waste: leaderboard

Metric: Success-conditioned safe rate (%): the share of agent trajectories that meet the case's own safety criterion at every planning and tool-calling step (a benign case is safe when completed with the needed safeguards or paused for clarification, an adversarial case when the agent refuses or halts before harm), among trajectories labeled safe or unsafe by a DeepSeek-R1 evaluator (execution failures excluded); on the 67 benign user-induced cases of risk type IP (plans that waste resources or cost (for example failing to parallelize)); LPS-Bench's 570 human-reviewed long-horizon tool-use cases with simulated MCP-style toolkits, each model in one LangChain agent at temperature 1 with up to 100 steps; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.595.52#138
2Gemini 3 Pro94.03#77
3GPT-5.189.55#131
4GPT-561.19#91
5Claude 3.5 Sonnet56.72#337
6Claude Sonnet 437.31#194
7Gemini 2.5 Pro31.34#145
8DeepSeek V3.128.36#260
9DeepSeek V3.214.93#198
10Qwen 3 32B10.45#424
11Llama 3.1 8B Instruct4.48#1018
12Llama 3.1 70B Instruct4.48#548
13Qwen 3 8B2.99#667

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=lps-bench-inefficient-planning-and-resource-waste · How It Works · Data refreshed daily, snapshot 2026-10-11.