LPS-Bench - Benign User Risks: leaderboard
Metric: Success-conditioned safe rate (%): the share of agent trajectories that meet the case's own safety criterion at every planning and tool-calling step (a benign case is safe when completed with the needed safeguards or paused for clarification, an adversarial case when the agent refuses or halts before harm), among trajectories labeled safe or unsafe by a DeepSeek-R1 evaluator (execution failures excluded); unweighted mean of the four benign user-induced risk types (dependency and ordering hazards, rigid over-compliance, ambiguity-induced false assumptions, inefficient planning; 252 cases); LPS-Bench's 570 human-reviewed long-horizon tool-use cases with simulated MCP-style toolkits, each model in one LangChain agent at temperature 1 with up to 100 steps; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Claude Sonnet 4.5 | 58.55 | #138 |
| 2 | Gemini 3 Pro | 55.35 | #77 |
| 3 | GPT-5.1 | 49.95 | #131 |
| 4 | GPT-5 | 38.29 | #91 |
| 5 | Claude Sonnet 4 | 35.33 | #194 |
| 6 | Claude 3.5 Sonnet | 30.9 | #337 |
| 7 | Gemini 2.5 Pro | 27.62 | #145 |
| 8 | DeepSeek V3.2 | 27.17 | #198 |
| 9 | DeepSeek V3.1 | 23.85 | #260 |
| 10 | Qwen 3 32B | 9.86 | #424 |
| 11 | Qwen 3 8B | 7.9 | #667 |
| 12 | Llama 3.1 70B Instruct | 4.5 | #548 |
| 13 | Llama 3.1 8B Instruct | 4.01 | #1018 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=lps-bench-benign-user-risks · How It Works · Data refreshed daily, snapshot 2026-10-11.