DUMA-Bench (Solo): leaderboard
Metric: Attack success rate (%; share of trials in which the agent violates a task's executable security assertions, plus GPT-4o-judged communication assertions on 9 of the 35 tau2-bench-derived tasks in 8 vulnerability domains; repeated independent runs of each task; passive-user regime: the agent acts on the environment without an active user). Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 5 |
| 2 | Claude Opus 4.5 | 7.1 |
| 3 | Claude Haiku 4.5 | 7.9 |
| 4 | GPT-5 | 25.7 |
| 5 | GPT-5 Mini | 27.9 |
| 6 | GLM-4.7 | 27.9 |
| 7 | GPT-4.1 Mini | 29.3 |
| 8 | GPT-4o | 31.4 |
| 9 | GPT-5 Nano | 31.4 |
| 10 | Qwen 3.5 Flash | 32.9 |
| 11 | GPT-4.1 | 33.6 |
| 12 | DeepSeek V3.2 | 34.3 |
| 13 | Qwen 3.5 Plus | 39.3 |
| 14 | GPT-4o Mini | 44 |
Interactive version: theaggregate.ai/benchmark?slug=duma-bench-solo · How It Works · Data refreshed daily, snapshot 2026-09-26.