TrustShiftProbe - Location Navigation: leaderboard
Metric: Attack success rate (%, lower is better; share of the location navigation attack sessions whose switched payload succeeds; each model as a ReAct agent connected to one MCP server that behaves benignly during a conditioning phase and then switches to an adversarial payload; 90 cases per domain; temperature 0, 15 tool calls; no defense). Source: arxiv.org. Saturation forecast: Around February 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.8 | 55.3 |
| 2 | O4 Mini (2025-04-16) | 61 |
| 3 | GPT-5 | 61.6 |
| 4 | Qwen 3.5 Flash | 61.8 |
| 5 | GPT-4.1 | 66.3 |
| 6 | Grok 4.3 | 76.5 |
Interactive version: theaggregate.ai/benchmark?slug=trustshiftprobe-location-navigation · How It Works · Data refreshed daily, snapshot 2026-09-29.