DuMateBench: leaderboard
Metric: Final score (%; 0.3 x partial pass rate + 0.7 x Gemini-3.1-Pro-Preview judge score of the produced artifacts against task-specific rubrics; 200 tasks reconstructed from anonymized, privacy-screened DuMate production sessions, each in an isolated Docker container with insufficient, unstable and noisy environment conditions; each agent keeps its native control loop and runtime settings). Source: arxiv.org. Saturation forecast: Around December 2026. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V4 Pro | 82.23 |
| 2 | Claude Opus 4.8 | 81.81 |
| 3 | GPT-5.5 | 80.44 |
| 4 | GLM-5.2 | 77.61 |
Interactive version: theaggregate.ai/benchmark?slug=dumatebench · How It Works · Data refreshed daily, snapshot 2026-09-29.