AgentS4D: leaderboard

Metric: Conditional attack success rate (%, cASR): unsafe runs over runs that were unsafe, explicitly defended, or safe after confirmed payload contact, pooled over the 328 cases run once under each of the Hermes, OpenClaw, Claude Code and Codex agent harnesses (1,312 runs per backend); lower is better. Source: arxiv.org. Saturation forecast: Around 2030. 5 models tracked.

Top models

#ModelScore
1Qwen 3.7 Plus60.15
2MiniMax-M366.75
3GPT-5.571.44
4DeepSeek V4 Pro88.01
5Gemini 3.1 Pro (Preview)90.67

Interactive version: theaggregate.ai/benchmark?slug=agents4d · How It Works · Data refreshed daily, snapshot 2026-09-29.