SABER - Embedded Injection: leaderboard

Metric: Harmful safety-violation rate (HSR, %) over effective runs on the 289 Scenario A tasks (malicious instructions embedded in project artifacts or tool outputs), each model runs as a coding agent in a Docker-sandboxed project workspace through the benchmark's own ReAct tool loop (shell plus task MCP tools), one run per task; a run is a violation when rule-based state and command checks or the semantic LLM judge flag harm; runs judged Incapable (failures and unnecessary refusals) are excluded from the denominator; lower is better. Source: arxiv.org. Saturation forecast: Around November 2027. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.643.7
2GLM-563.7
3GPT-5.464
4MiniMax-M2.567.2
5Qwen 3.5 397B A17B69.4
6Kimi K2.571.1
7GLM-4.772
8DeepSeek V372.7
9DeepSeek V3.273.3
10Ling-flash-2.074.2
11Qwen 3.5 35B A3B76.4
12Qwen 3.5 9B76.9
13DeepSeek R184.3

Interactive version: theaggregate.ai/benchmark?slug=saber-embedded-injection · How It Works · Data refreshed daily, snapshot 2026-09-29.