SABER - Contextual Warnings: leaderboard

Metric: Harmful safety-violation rate (HSR, %) over effective runs on the 241 Scenario C tasks (workspace evidence makes direct execution unsafe), each model runs as a coding agent in a Docker-sandboxed project workspace through the benchmark's own ReAct tool loop (shell plus task MCP tools), one run per task; a run is a violation when rule-based state and command checks or the semantic LLM judge flag harm; runs judged Incapable (failures and unnecessary refusals) are excluded from the denominator; lower is better. Source: arxiv.org. Saturation forecast: Around 2028. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.663.1
2GPT-5.466.5
3DeepSeek V378.3
4Ling-flash-2.081.3
5Qwen 3.5 9B83.2
6GLM-583.4
7Qwen 3.5 397B A17B85
8Kimi K2.585.5
9Qwen 3.5 35B A3B85.5
10GLM-4.785.6
11MiniMax-M2.587.8
12DeepSeek V3.290.2
13DeepSeek R191.9

Interactive version: theaggregate.ai/benchmark?slug=saber-contextual-warnings · How It Works · Data refreshed daily, snapshot 2026-09-29.