TrustShiftProbe: leaderboard
Metric: Attack success rate (%, lower is better; share of the 360 attack sessions in four domains whose switched payload succeeds (wrong or unavailable final answer, or exfiltration to a simulated sink); each model as a ReAct agent connected to one MCP server that behaves benignly during a conditioning phase and then switches to an adversarial payload; 90 cases per domain; temperature 0, 15 tool calls; no defense). Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 60.2 |
| 2 | O4 Mini (2025-04-16) | 69.2 |
| 3 | Qwen 3.5 Flash | 69.5 |
| 4 | Claude Opus 4.8 | 70.9 |
| 5 | Grok 4.3 | 73.8 |
| 6 | GPT-4.1 | 74.1 |
Interactive version: theaggregate.ai/benchmark?slug=trustshiftprobe · How It Works · Data refreshed daily, snapshot 2026-09-29.