DuMateBench: leaderboard

Metric: Final score (%; 0.3 x partial pass rate + 0.7 x Gemini-3.1-Pro-Preview judge score of the produced artifacts against task-specific rubrics; 200 tasks reconstructed from anonymized, privacy-screened DuMate production sessions, each in an isolated Docker container with insufficient, unstable and noisy environment conditions; each agent keeps its native control loop and runtime settings). Source: arxiv.org. Saturation forecast: Around December 2026. 20 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro82.23
2Claude Opus 4.881.81
3GPT-5.580.44
4GLM-5.277.61

Interactive version: theaggregate.ai/benchmark?slug=dumatebench · How It Works · Data refreshed daily, snapshot 2026-09-29.