TrustShiftProbe: leaderboard

Metric: Attack success rate (%, lower is better; share of the 360 attack sessions in four domains whose switched payload succeeds (wrong or unavailable final answer, or exfiltration to a simulated sink); each model as a ReAct agent connected to one MCP server that behaves benignly during a conditioning phase and then switches to an adversarial payload; 90 cases per domain; temperature 0, 15 tool calls; no defense). Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.

Top models

#ModelScore
1GPT-560.2
2O4 Mini (2025-04-16)69.2
3Qwen 3.5 Flash69.5
4Claude Opus 4.870.9
5Grok 4.373.8
6GPT-4.174.1

Interactive version: theaggregate.ai/benchmark?slug=trustshiftprobe · How It Works · Data refreshed daily, snapshot 2026-09-29.