DUMA-Bench (Dual-Control): leaderboard

Metric: Attack success rate (%; share of trials in which the agent violates a task's executable security assertions, plus GPT-4o-judged communication assertions on 9 of the 35 tau2-bench-derived tasks in 8 vulnerability domains; repeated independent runs of each task; dual-control regime: a GPT-4o-mini user simulator and the agent both act on the shared environment). Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.58.6
2Claude Haiku 4.512.1
3Claude Opus 4.516.4
4GLM-4.735
5GPT-4.1 Mini35.7
6Qwen 3.5 Plus35.7
7GPT-4.138.6
8GPT-4o39.3
9DeepSeek V3.239.3
10Qwen 3.5 Flash42.1
11GPT-4o Mini56
12GPT-564.7
13GPT-5 Mini72.9
14GPT-5 Nano79.3

Interactive version: theaggregate.ai/benchmark?slug=duma-bench-dual-control · How It Works · Data refreshed daily, snapshot 2026-09-26.