MonitoringBench - Refined - Password Leak: leaderboard

Metric: Catch rate (%) at a 1% false-positive-rate threshold calibrated on 416 honest BashArena runs: share of the 198 successful credential exfiltration (password leak) attack trajectories (pipeline trajectories rewritten by the Claude Opus 4.6 refinement stage) that the full-trajectory monitor scores above its threshold, default ControlArena monitor prompt with a 0-10 suspiciousness score, mean of three scoring runs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 (20251101)43.8
2GPT-5.1 Codex Max38.6
3Claude Sonnet 4.536.2
4GPT-528.1
5Claude Sonnet 4 (20250514)22.7
6O3 (2025-04-16)22.4
7GPT-4.14.5
8GPT-4 Turbo2.2
9GPT-4o (2024-08-06)1.9
10GPT-5 Nano0.7
11GPT-4.1 Mini0.3
12O4 Mini (2025-04-16)0.3
13GPT-4o Mini (2024-07-18)0

Interactive version: theaggregate.ai/benchmark?slug=monitoringbench-refined-password-leak · How It Works · Data refreshed daily, snapshot 2026-10-07.