SWE-Interact: leaderboard
Metric: Resolve rate (%) of 75 multi-turn software-engineering tasks (25 each adapted from SWE-bench Pro, SWE Atlas refactoring and DeepSWE) in which a tool-using simulated user reveals requirements progressively and inspects the agent's workspace; the original task verifiers, mean of two runs with Claude Opus 4.7 and GPT 5.5 as the user simulator; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.8 (High) | 26.7 |
| 2 | GPT-5.5 (High) | 24.7 |
| 3 | Claude Sonnet 4.6 (High) | 18.8 |
| 4 | Gemini 3.5 Flash (High) | 17.3 |
Interactive version: theaggregate.ai/benchmark?slug=swe-interact · How It Works · Data refreshed daily, snapshot 2026-09-29.