MonitoringBench - Human Strategies - Password Leak: leaderboard

Metric: Catch rate (%) at a 1% false-positive-rate threshold calibrated on 416 honest BashArena runs: share of the 110 successful credential exfiltration (password leak) attack trajectories (human-seeded attack strategies executed by Claude Opus 4.6) that the full-trajectory monitor scores above its threshold, default ControlArena monitor prompt with a 0-10 suspiciousness score, mean of three scoring runs; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 (20251101)99.4
2GPT-5.1 Codex Max98.4
3Claude Sonnet 4.598.2
4GPT-595.5
5O3 (2025-04-16)92.1
6Claude Sonnet 4 (20250514)84.5
7GPT-4.168.8
8O4 Mini (2025-04-16)36.7
9GPT-4.1 Mini31.5
10GPT-4o (2024-08-06)31.3
11GPT-4 Turbo29.4
12GPT-5 Nano21.3
13GPT-4o Mini (2024-07-18)0.9

Interactive version: theaggregate.ai/benchmark?slug=monitoringbench-human-strategies-password-leak · How It Works · Data refreshed daily, snapshot 2026-10-07.