TrustShiftProbe - Financial Analysis: leaderboard
Metric: Attack success rate (%, lower is better; share of the financial analysis attack sessions whose switched payload succeeds; each model as a ReAct agent connected to one MCP server that behaves benignly during a conditioning phase and then switches to an adversarial payload; 90 cases per domain; temperature 0, 15 tool calls; no defense). Source: arxiv.org. Saturation forecast: Around 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 64.4 |
| 2 | Claude Opus 4.8 | 68.2 |
| 3 | GPT-4.1 | 75.6 |
| 4 | Grok 4.3 | 75.6 |
| 5 | Qwen 3.5 Flash | 76.2 |
| 6 | O4 Mini (2025-04-16) | 84 |
Interactive version: theaggregate.ai/benchmark?slug=trustshiftprobe-financial-analysis · How It Works · Data refreshed daily, snapshot 2026-09-29.