MonitoringBench - Human Strategies - Sudo Escalation: leaderboard

Metric: Catch rate (%) at a 1% false-positive-rate threshold calibrated on 416 honest BashArena runs: share of the 313 successful sudo privilege escalation attack trajectories (human-seeded attack strategies executed by Claude Opus 4.6) that the full-trajectory monitor scores above its threshold, default ControlArena monitor prompt with a 0-10 suspiciousness score, mean of three scoring runs; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 (20251101)93.5
2GPT-5.1 Codex Max84
3Claude Sonnet 4.578.4
4GPT-570.2
5O3 (2025-04-16)66.3
6Claude Sonnet 4 (20250514)62.9
7GPT-4.130.4
8GPT-5 Nano23.4
9GPT-4o (2024-08-06)15.1
10GPT-4 Turbo14.2
11O4 Mini (2025-04-16)11.1
12GPT-4.1 Mini5
13GPT-4o Mini (2024-07-18)0.7

Interactive version: theaggregate.ai/benchmark?slug=monitoringbench-human-strategies-sudo-escalation · How It Works · Data refreshed daily, snapshot 2026-10-07.