DuelLab Overall — leaderboard

DuelLab evaluates model-generated game-playing programs by compiling submitted code and running head-to-head tournaments on hidden abstract strategy games.

Metric: Avg score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 51 models tracked.

Top models

#ModelScore
1Claude Opus 4.774.4
2GPT-5.474.2
3GPT-5.566.5
4GPT-5.266.1
5Claude Opus 4.665.2
6DeepSeek V4 Pro65
7GLM-5.163.1
8GPT-5.3 Codex62.4
9Gemini 3.1 Pro (Preview)62.3
10Claude Sonnet 4.660.9
11Qwen 3.6 Plus58.6
12GPT-5.4 Nano55.6
13GLM-555.5
14MiMo-V2.5-Pro54.2
15DeepSeek V4 Flash52.5

Interactive version: theaggregate.ai/benchmark?slug=duellab-overall · How the rankings work · Data refreshed daily, snapshot 2026-07-22.