DUMA-Bench (Solo): leaderboard

Metric: Attack success rate (%; share of trials in which the agent violates a task's executable security assertions, plus GPT-4o-judged communication assertions on 9 of the 35 tau2-bench-derived tasks in 8 vulnerability domains; repeated independent runs of each task; passive-user regime: the agent acts on the environment without an active user). Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.55
2Claude Opus 4.57.1
3Claude Haiku 4.57.9
4GPT-525.7
5GPT-5 Mini27.9
6GLM-4.727.9
7GPT-4.1 Mini29.3
8GPT-4o31.4
9GPT-5 Nano31.4
10Qwen 3.5 Flash32.9
11GPT-4.133.6
12DeepSeek V3.234.3
13Qwen 3.5 Plus39.3
14GPT-4o Mini44

Interactive version: theaggregate.ai/benchmark?slug=duma-bench-solo · How It Works · Data refreshed daily, snapshot 2026-09-26.