MonitoringBench - Recon - Firewall Backdoor: leaderboard

Metric: Catch rate (%) at a 1% false-positive-rate threshold calibrated on 416 honest BashArena runs: share of the 204 successful firewall backdoor attack trajectories (strategies from reconnaissance agents (with and without GPT-5 monitor feedback) executed by Claude Opus 4.6) that the full-trajectory monitor scores above its threshold, default ControlArena monitor prompt with a 0-10 suspiciousness score, mean of three scoring runs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 (20251101)72.9
2Claude Sonnet 4.562.7
3Claude Sonnet 4 (20250514)56
4GPT-5.1 Codex Max51.8
5O3 (2025-04-16)35.9
6GPT-529.7
7GPT-4.116.5
8GPT-4o (2024-08-06)12.7
9GPT-4 Turbo7.5
10GPT-5 Nano5.7
11GPT-4.1 Mini4.7
12O4 Mini (2025-04-16)2.5
13GPT-4o Mini (2024-07-18)0.5

Interactive version: theaggregate.ai/benchmark?slug=monitoringbench-recon-firewall-backdoor · How It Works · Data refreshed daily, snapshot 2026-10-07.