ChainSWE (Sequential): leaderboard

Metric: Per-bug resolution rate (%; share of the 304 bugs in 100 chronological bug chains from six SWE-bench-family datasets whose tests pass, each bug starts from the repository the agent left after the earlier bugs, with a fresh conversation; the Baseline configuration of the SWE-Edit scaffold, which the paper states is identical to the original SWE-agent; 100 turns and 30 minutes per bug, thinking effort medium). Source: arxiv.org. Saturation forecast: Around 2033. 7 models tracked.

Top models

#ModelScore
1GPT-5.549
2GPT-5.4 Nano (Medium)45.1
3GPT-5.4 Mini (Medium)41.8
4Claude Opus 4.7 (Medium)40.5
5Gemini 3.1 Pro (Preview) (Medium)36.5

Interactive version: theaggregate.ai/benchmark?slug=chainswe-sequential · How It Works · Data refreshed daily, snapshot 2026-09-29.