SLBench - Unsafe Rate: leaderboard
Metric: Unsafe outcome rate: share of runs in which a case-specific violation signal fired (%) of the 86 audited SLBench cases (39 controls, 47 violation cases) built from logical relations between clauses of real agent skill files, graded deterministically from repository state, artifacts and command traces with unsafe-first precedence; each backbone runs in its vendor agent harness (Codex CLI or Claude Code); lower is better. Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.7 | 35.1 |
| 2 | GPT-5.3 Codex | 39.5 |
| 3 | Claude Sonnet 4.6 | 43 |
| 4 | Claude Haiku 4.5 | 53.5 |
| 5 | GPT-5.4 Mini | 57 |
| 6 | GPT-5.5 | 70.2 |
Interactive version: theaggregate.ai/benchmark?slug=slbench-unsafe-rate · How It Works · Data refreshed daily, snapshot 2026-09-29.