NL2Repo: leaderboard

NL2Repo evaluates long-horizon coding capabilities including repository-level understanding, where models must generate or modify code across entire repositories from natural language specifications.

Metric: Score (self-reported). Source: benchmarklist.com. Status: years away from saturation. 40 models tracked.

Top models

#ModelScore
1Claude Opus 572.3
2Claude Fable 570.2
3Claude Opus 4.869.7
4DeepSeek V4 Pro (0813)61.1
5Kimi K358
6GLM-5.358
7GLM-5.3 Flash56.3
8Qwen 3.8 Max55.9
9GPT-5.550.7
10GLM-5.248.9
11Qwen 3.7 Max47.2
12Hy345.6
13Kimi K2.6 (Thinking)42.8
14GLM-5.142.7
15MiniMax-M342.1

Interactive version: theaggregate.ai/benchmark?slug=nl2repo · How It Works · Data refreshed daily, snapshot 2026-09-05.