LPS-Bench - Ambiguity-induced False Assumptions: leaderboard

Metric: Success-conditioned safe rate (%): the share of agent trajectories that meet the case's own safety criterion at every planning and tool-calling step (a benign case is safe when completed with the needed safeguards or paused for clarification, an adversarial case when the agent refuses or halts before harm), among trajectories labeled safe or unsafe by a DeepSeek-R1 evaluator (execution failures excluded); on the 68 benign user-induced cases of risk type FA (undefined details the agent should clarify rather than assume); LPS-Bench's 570 human-reviewed long-horizon tool-use cases with simulated MCP-style toolkits, each model in one LangChain agent at temperature 1 with up to 100 steps; higher is better. Source: arxiv.org. 13 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.126.47#131
2Gemini 3 Pro16.18#77
3GPT-514.71#91
4DeepSeek V3.111.76#260
5Claude Sonnet 4.55.88#138
6DeepSeek V3.24.41#198
7Qwen 3 8B4.41#667
8Qwen 3 32B2.94#424
9Claude Sonnet 42.94#194
10Claude 3.5 Sonnet2.94#337
11Gemini 2.5 Pro1.47#145
12Llama 3.1 8B Instruct1.47#1018
13Llama 3.1 70B Instruct0#548

Interactive version: theaggregate.ai/benchmark?slug=lps-bench-ambiguity-induced-false-assumptions · How It Works · Data refreshed daily, snapshot 2026-10-11.