LPS-Bench - Prompt Injection and Jailbreaks: leaderboard

Metric: Success-conditioned safe rate (%): the share of agent trajectories that meet the case's own safety criterion at every planning and tool-calling step (a benign case is safe when completed with the needed safeguards or paused for clarification, an adversarial case when the agent refuses or halts before harm), among trajectories labeled safe or unsafe by a DeepSeek-R1 evaluator (execution failures excluded); on the 67 adversarial user-induced cases of risk type PI (crafted inputs that override the planning logic or safety constraints); LPS-Bench's 570 human-reviewed long-horizon tool-use cases with simulated MCP-style toolkits, each model in one LangChain agent at temperature 1 with up to 100 steps; higher is better. Source: arxiv.org. 13 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 498.51#194
2GPT-5.198.51#131
3Claude Sonnet 4.594.03#138
4Gemini 3 Pro79.1#77
5Claude 3.5 Sonnet53.73#337
6GPT-537.31#91
7Gemini 2.5 Pro22.39#145
8Llama 3.1 8B Instruct5.97#1018
9Llama 3.1 70B Instruct5.97#548
10Qwen 3 8B2.99#667
11Qwen 3 32B1.49#424
12DeepSeek V3.11.49#260
13DeepSeek V3.20#198

Interactive version: theaggregate.ai/benchmark?slug=lps-bench-prompt-injection-and-jailbreaks · How It Works · Data refreshed daily, snapshot 2026-10-11.