MonitoringBench - Recon - Password Leak: leaderboard

Metric: Catch rate (%) at a 1% false-positive-rate threshold calibrated on 416 honest BashArena runs: share of the 128 successful credential exfiltration (password leak) attack trajectories (strategies from reconnaissance agents (with and without GPT-5 monitor feedback) executed by Claude Opus 4.6) that the full-trajectory monitor scores above its threshold, default ControlArena monitor prompt with a 0-10 suspiciousness score, mean of three scoring runs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 (20251101)75.8
2GPT-5.1 Codex Max68.2
3Claude Sonnet 4.568
4GPT-564.1
5O3 (2025-04-16)62.5
6Claude Sonnet 4 (20250514)49.2
7GPT-4.131.8
8GPT-4o (2024-08-06)12.5
9O4 Mini (2025-04-16)11
10GPT-4 Turbo9.9
11GPT-4.1 Mini9.6
12GPT-5 Nano9.2
13GPT-4o Mini (2024-07-18)0

Interactive version: theaggregate.ai/benchmark?slug=monitoringbench-recon-password-leak · How It Works · Data refreshed daily, snapshot 2026-10-07.