MT-AgentRisk - Notion: leaderboard
Metric: Attack success rate (%): share of the 15 Notion-MCP tasks the agent completes, on MT-AgentRisk multi-turn attack sequences (365 harmful tool-use tasks from OpenAgentSafety, SafeArena, P2SQL and harmful-converted MCPMark Notion tasks, each split by a taxonomy of addition and decomposition attacks into 2 to 7 innocuous-looking turns), each model as an OpenHands agent with real MCP tools at default settings and no defense; a GPT-4.1 judge labels each trajectory complete, rejected or failed; lower is better. Source: arxiv.org. 6 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Claude Sonnet 4.5 | 26.67 | #138 |
| 2 | Seed-1.6 | 60 | #257 |
| 3 | Qwen3 Coder | 60 | #285 |
| 4 | DeepSeek V3.2 | 73.33 | #198 |
| 5 | Gemini 3 Flash | 80 | #93 |
| 6 | GPT-5.2 (Medium) | 86.67 | #105 (GPT-5.2) |
Interactive version: theaggregate.ai/benchmark?slug=mt-agentrisk-notion · How It Works · Data refreshed daily, snapshot 2026-10-11.