NL2Repo — leaderboard

NL2Repo evaluates long-horizon coding capabilities including repository-level understanding, where models must generate or modify code across entire repositories from natural language specifications.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 12 models tracked.

Top models

#ModelScore
1Claude Opus 4.869.7
2GPT-5.550.7
3GLM-5.248.9
4Claude Opus 4.6 (Max)47.6
5Qwen 3.7 Max (Max)47.2
6GLM-5.142.7
7MiniMax-M342.1
8Qwen 3.7 Plus41.1
9DeepSeek V4 Pro35.5
10Qwen 3.6 Plus34.4
11Gemini 3.1 Pro (Preview)33.4

Interactive version: theaggregate.ai/benchmark?slug=nl2repo · How the rankings work · Data refreshed daily, snapshot 2026-07-22.