Django Benchmark Suite - Method Completion: leaderboard

Metric: pass@1 (%; Method completion: the model completes a method from its first half and the surrounding file, without a docstring; the output is inserted into the file and counts only when every Django test mapped to the method passes; 359 held-out methods of a leakage-controlled Django repository snapshot; greedy decoding, no agent scaffold). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Gemini 3 Pro78.27
2GPT-577.99
3Claude Sonnet 4.577.44
4Claude Opus 4.576.6
5Gemini 3 Flash73.26
6GPT-4.164.9
7Qwen 2.5 Coder 32B Instruct60.17
8Claude Haiku 4.558.77
9Gemini 2.5 Pro51.25
10Qwen 2.5 Coder 7B Instruct50.97
11Qwen 3 32B43.45

Interactive version: theaggregate.ai/benchmark?slug=django-benchmark-suite-method-completion · How It Works · Data refreshed daily, snapshot 2026-09-29.