Multi-Docker-Eval — leaderboard
Tests LLMs on constructing executable Docker environments for real-world software repositories. Given a GitHub repo, the model must produce a working Dockerfile and configuration.
Metric: Resolved (%). Source: github.com. Status: saturation imminent. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Kimi K2.5 | 41.82 |
| 2 | GLM-4.7 | 40.62 |
| 3 | DeepSeek V3.1 | 37.72 |
| 4 | Kimi K2 0905 | 37.62 |
| 5 | DeepSeek V3.2 | 36.63 |
| 6 | Kimi K2 (Thinking) | 36.53 |
| 7 | Claude Sonnet 4 | 35.53 |
| 8 | Kimi K2 (0711) | 34.23 |
| 9 | GPT-5 Mini | 34.13 |
| 10 | MiMo-V2-Flash | 32.83 |
| 11 | Gemini 2.5 Flash | 29.44 |
| 12 | GPT-OSS-120B | 27.35 |
| 13 | DeepSeek R1 | 26.65 |
| 14 | Qwen 3 235B A22B | 23.65 |
| 15 | GPT-OSS-20B | 17.17 |
Interactive version: theaggregate.ai/benchmark?slug=multi-docker-eval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.