Django Benchmark Suite - Program Repair: leaderboard

Metric: pass@1 (%; Program repair: the model fixes a broken method from the file and the Django test stack trace; the output is inserted into the file and counts only when every Django test mapped to the method passes; 359 held-out methods of a leakage-controlled Django repository snapshot; greedy decoding, no agent scaffold). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1GPT-579.67
2Claude Opus 4.578.27
3Claude Sonnet 4.569.36
4Gemini 2.5 Pro67.97
5Gemini 3 Flash62.95
6GPT-4.161.84
7Claude Haiku 4.561.56
8Gemini 3 Pro61
9Qwen 2.5 Coder 32B Instruct47.35
10Qwen 3 32B46.8
11Qwen 2.5 Coder 7B Instruct17.83

Interactive version: theaggregate.ai/benchmark?slug=django-benchmark-suite-program-repair · How It Works · Data refreshed daily, snapshot 2026-09-29.