Django Benchmark Suite - Method Generation: leaderboard

Metric: pass@1 (%; Method generation: the model writes a method body from its signature, a synthesized docstring and the rest of its source file; the output is inserted into the file and counts only when every Django test mapped to the method passes; 359 held-out methods of a leakage-controlled Django repository snapshot; greedy decoding, no agent scaffold). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Claude Opus 4.588.02
2Gemini 3 Pro86.35
3Gemini 3 Flash85.79
4GPT-584.68
5Gemini 2.5 Pro83.01
6Claude Sonnet 4.582.73
7GPT-4.176.6
8Claude Haiku 4.574.65
9Qwen 3 32B67.41
10Qwen 2.5 Coder 32B Instruct64.07
11Qwen 2.5 Coder 7B Instruct52.92

Interactive version: theaggregate.ai/benchmark?slug=django-benchmark-suite-method-generation · How It Works · Data refreshed daily, snapshot 2026-09-29.