SABER: leaderboard

Metric: Harmful safety-violation rate (HSR, %) over effective runs on all 716 tasks, each model runs as a coding agent in a Docker-sandboxed project workspace through the benchmark's own ReAct tool loop (shell plus task MCP tools), one run per task; a run is a violation when rule-based state and command checks or the semantic LLM judge flag harm; runs judged Incapable (failures and unnecessary refusals) are excluded from the denominator; lower is better. Source: arxiv.org. Saturation forecast: Around 2028. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.654.7
2GPT-5.463.9
3GLM-571
4DeepSeek V372.4
5Qwen 3.5 397B A17B73.4
6MiniMax-M2.573.7
7Ling-flash-2.075.4
8Kimi K2.576.1
9GLM-4.777
10Qwen 3.5 35B A3B77.3
11Qwen 3.5 9B78.6
12DeepSeek V3.279.6
13DeepSeek R184.7

Interactive version: theaggregate.ai/benchmark?slug=saber · How It Works · Data refreshed daily, snapshot 2026-09-29.