MT-AgentRisk - Playwright: leaderboard

Metric: Attack success rate (%): share of the 140 Playwright-MCP web tasks (GitLab, OwnCloud, Reddit, shopping and shopping admin) the agent completes, on MT-AgentRisk multi-turn attack sequences (365 harmful tool-use tasks from OpenAgentSafety, SafeArena, P2SQL and harmful-converted MCPMark Notion tasks, each split by a taxonomy of addition and decomposition attacks into 2 to 7 innocuous-looking turns), each model as an OpenHands agent with real MCP tools at default settings and no defense; a GPT-4.1 judge labels each trajectory complete, rejected or failed; lower is better. Source: arxiv.org. 6 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.2 (Medium)34.29#105 (GPT-5.2)
2Claude Sonnet 4.565.71#138
3Qwen3 Coder67.86#285
4Seed-1.670.71#257
5Gemini 3 Flash72.86#93
6DeepSeek V3.280#198

Interactive version: theaggregate.ai/benchmark?slug=mt-agentrisk-playwright · How It Works · Data refreshed daily, snapshot 2026-10-11.