SWE-Explore - Precision: leaderboard

Metric: Line-level precision (%; share of returned lines inside the core context) on 848 repository-exploration instances (issues from 203 repositories in 10 languages): given the repository and issue, the explorer returns five ranked code regions, scored against line-level core context distilled from independent successful repair trajectories; the Mini-SWE-Agent scaffold with each LLM; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 6 models tracked.

Top models

#ModelScore
1GPT-5.454.2
2Claude Sonnet 4.551.9
3GPT-5.4 Mini50.9
4Kimi K2.647.5
5Gemini 3 Pro42
6GLM-4.741.4

Interactive version: theaggregate.ai/benchmark?slug=swe-explore-precision · How It Works · Data refreshed daily, snapshot 2026-09-29.