EvoCode-Bench - Task Completion: leaderboard

Metric: Full-task completion (Comp, %): share of the 26 tasks in which at least one of four attempts passes the cumulative verifier at every round through the final one (fail-stop), over the 26 stateful multi-round coding tasks of EvoCode-Bench (227 rounds, 5 to 15 per task, requirements that extend, correct or contradict earlier ones, cumulative executable verifiers), each model driving the Terminus-2 agent harness in a Harbor Docker workspace; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.7 (High)42.3
2GPT-5.5 (High)38.5
3Claude Opus 4.634.6
4Kimi K2.623.1
5DeepSeek V4 Pro (High)19.2
6GLM-5.115.4
7Qwen 3.6 Plus (Thinking)15.4
8Gemini 3.1 Pro (Preview) (High)11.5
9Qwen 3.5 397B A17B0
10MiniMax-M2.70
11Seed 2.0 Pro (High)0
12DeepSeek V4 Flash (High)0

Interactive version: theaggregate.ai/benchmark?slug=evocode-bench-task-completion · How It Works · Data refreshed daily, snapshot 2026-10-07.