Multi-Docker-Eval — leaderboard

Tests LLMs on constructing executable Docker environments for real-world software repositories. Given a GitHub repo, the model must produce a working Dockerfile and configuration.

Metric: Resolved (%). Source: github.com. Status: saturation imminent. 16 models tracked.

Top models

#ModelScore
1Kimi K2.541.82
2GLM-4.740.62
3DeepSeek V3.137.72
4Kimi K2 090537.62
5DeepSeek V3.236.63
6Kimi K2 (Thinking)36.53
7Claude Sonnet 435.53
8Kimi K2 (0711)34.23
9GPT-5 Mini34.13
10MiMo-V2-Flash32.83
11Gemini 2.5 Flash29.44
12GPT-OSS-120B27.35
13DeepSeek R126.65
14Qwen 3 235B A22B23.65
15GPT-OSS-20B17.17

Interactive version: theaggregate.ai/benchmark?slug=multi-docker-eval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.