TrustShiftProbe - Browser Automation: leaderboard
Metric: Attack success rate (%, lower is better; share of the browser automation attack sessions whose switched payload succeeds; each model as a ReAct agent connected to one MCP server that behaves benignly during a conditioning phase and then switches to an adversarial payload; 90 cases per domain; temperature 0, 15 tool calls; no defense). Source: arxiv.org. Saturation forecast: Around 2035. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4.3 | 62.3 |
| 2 | Qwen 3.5 Flash | 62.8 |
| 3 | GPT-5 | 65.4 |
| 4 | O4 Mini (2025-04-16) | 72 |
| 5 | GPT-4.1 | 76.5 |
| 6 | Claude Opus 4.8 | 85.1 |
Interactive version: theaggregate.ai/benchmark?slug=trustshiftprobe-browser-automation · How It Works · Data refreshed daily, snapshot 2026-09-29.