DUMA-Bench (Dual-Control): leaderboard
Metric: Attack success rate (%; share of trials in which the agent violates a task's executable security assertions, plus GPT-4o-judged communication assertions on 9 of the 35 tau2-bench-derived tasks in 8 vulnerability domains; repeated independent runs of each task; dual-control regime: a GPT-4o-mini user simulator and the agent both act on the shared environment). Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 8.6 |
| 2 | Claude Haiku 4.5 | 12.1 |
| 3 | Claude Opus 4.5 | 16.4 |
| 4 | GLM-4.7 | 35 |
| 5 | GPT-4.1 Mini | 35.7 |
| 6 | Qwen 3.5 Plus | 35.7 |
| 7 | GPT-4.1 | 38.6 |
| 8 | GPT-4o | 39.3 |
| 9 | DeepSeek V3.2 | 39.3 |
| 10 | Qwen 3.5 Flash | 42.1 |
| 11 | GPT-4o Mini | 56 |
| 12 | GPT-5 | 64.7 |
| 13 | GPT-5 Mini | 72.9 |
| 14 | GPT-5 Nano | 79.3 |
Interactive version: theaggregate.ai/benchmark?slug=duma-bench-dual-control · How It Works · Data refreshed daily, snapshot 2026-09-26.