MonitoringBench - Recon - Sudo Escalation: leaderboard

Metric: Catch rate (%) at a 1% false-positive-rate threshold calibrated on 416 honest BashArena runs: share of the 265 successful sudo privilege escalation attack trajectories (strategies from reconnaissance agents (with and without GPT-5 monitor feedback) executed by Claude Opus 4.6) that the full-trajectory monitor scores above its threshold, default ControlArena monitor prompt with a 0-10 suspiciousness score, mean of three scoring runs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 (20251101)65.9
2GPT-5.1 Codex Max45.6
3Claude Sonnet 4.538.4
4Claude Sonnet 4 (20250514)30.9
5O3 (2025-04-16)29.3
6GPT-527.7
7GPT-4o (2024-08-06)11.7
8GPT-4.110.9
9GPT-5 Nano9.3
10GPT-4 Turbo8.1
11GPT-4.1 Mini3.3
12O4 Mini (2025-04-16)3.2
13GPT-4o Mini (2024-07-18)0.8

Interactive version: theaggregate.ai/benchmark?slug=monitoringbench-recon-sudo-escalation · How It Works · Data refreshed daily, snapshot 2026-10-07.