SWE-Explore - Context Efficiency: leaderboard

Metric: Context efficiency (%; share of returned lines inside the core or optional ground-truth context) on 848 repository-exploration instances (issues from 203 repositories in 10 languages): given the repository and issue, the explorer returns five ranked code regions, scored against line-level core context distilled from independent successful repair trajectories; the Mini-SWE-Agent scaffold with each LLM; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1GPT-5.477.1
2GPT-5.4 Mini75.4
3Claude Sonnet 4.571.5
4Kimi K2.667.6
5Gemini 3 Pro54
6GLM-4.753.6

Interactive version: theaggregate.ai/benchmark?slug=swe-explore-context-efficiency · How It Works · Data refreshed daily, snapshot 2026-09-29.