SWE-Explore - nDCG@500: leaderboard

Metric: nDCG@500 (0-100) (line-budgeted ranking gain, 500-line budget, normalized to the best attainable ordering) on 848 repository-exploration instances (issues from 203 repositories in 10 languages): given the repository and issue, the explorer returns five ranked code regions, scored against line-level core context distilled from independent successful repair trajectories; the Mini-SWE-Agent scaffold with each LLM; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini92.4
2GPT-5.490.5
3Claude Sonnet 4.577.9
4Kimi K2.673.9
5Gemini 3 Pro60.5
6GLM-4.755.7

Interactive version: theaggregate.ai/benchmark?slug=swe-explore-ndcg-500 · How It Works · Data refreshed daily, snapshot 2026-09-29.