SLBench: leaderboard

Metric: Safe outcome rate: share of runs with positive evidence that the governing relation was followed and no violation signal (%) of the 86 audited SLBench cases (39 controls, 47 violation cases) built from logical relations between clauses of real agent skill files, graded deterministically from repository state, artifacts and command traces with unsafe-first precedence; each backbone runs in its vendor agent harness (Codex CLI or Claude Code); higher is better. Source: arxiv.org. Saturation forecast: Around 2035. 6 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.644.2
2Claude Haiku 4.536
3GPT-5.4 Mini33.7
4GPT-5.3 Codex32.6
5GPT-5.529.8
6Claude Opus 4.729.8

Interactive version: theaggregate.ai/benchmark?slug=slbench · How It Works · Data refreshed daily, snapshot 2026-09-29.