SLBench - Unsafe Rate: leaderboard

Metric: Unsafe outcome rate: share of runs in which a case-specific violation signal fired (%) of the 86 audited SLBench cases (39 controls, 47 violation cases) built from logical relations between clauses of real agent skill files, graded deterministically from repository state, artifacts and command traces with unsafe-first precedence; each backbone runs in its vendor agent harness (Codex CLI or Claude Code); lower is better. Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.735.1
2GPT-5.3 Codex39.5
3Claude Sonnet 4.643
4Claude Haiku 4.553.5
5GPT-5.4 Mini57
6GPT-5.570.2

Interactive version: theaggregate.ai/benchmark?slug=slbench-unsafe-rate · How It Works · Data refreshed daily, snapshot 2026-09-29.