SABER - Risky Self-Selection: leaderboard

Metric: Harmful safety-violation rate (HSR, %) over effective runs on the 186 Scenario B tasks (benign requests where a risky operational shortcut is available), each model runs as a coding agent in a Docker-sandboxed project workspace through the benchmark's own ReAct tool loop (shell plus task MCP tools), one run per task; a run is a violation when rule-based state and command checks or the semantic LLM judge flag harm; runs judged Incapable (failures and unnecessary refusals) are excluded from the denominator; lower is better. Source: arxiv.org. Saturation forecast: Around 2030. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.660.2
2GPT-5.460.6
3DeepSeek V363.9
4Qwen 3.5 397B A17B64
5MiniMax-M2.565.2
6GLM-566.3
7Qwen 3.5 35B A3B67.3
8Ling-flash-2.069.3
9Kimi K2.571.8
10GLM-4.773.1
11DeepSeek V3.274.8
12Qwen 3.5 9B75
13DeepSeek R175.9

Interactive version: theaggregate.ai/benchmark?slug=saber-risky-self-selection · How It Works · Data refreshed daily, snapshot 2026-09-29.