ChainSWE (Oracle): leaderboard

Metric: Per-bug resolution rate (%; share of the 304 bugs in 100 chronological bug chains from six SWE-bench-family datasets whose tests pass, each bug starts from its chain with the earlier gold fixes applied; the Baseline configuration of the SWE-Edit scaffold, which the paper states is identical to the original SWE-agent; 100 turns and 30 minutes per bug, thinking effort medium). Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScore
1GPT-5.569.1
2GPT-5.4 Nano (Medium)64.8
3Claude Opus 4.7 (Medium)64.5
4GPT-5.4 Mini (Medium)61.8
5Gemini 3.1 Pro (Preview) (Medium)61.8

Interactive version: theaggregate.ai/benchmark?slug=chainswe-oracle · How It Works · Data refreshed daily, snapshot 2026-09-29.