SWE-Interact: leaderboard

Metric: Resolve rate (%) of 75 multi-turn software-engineering tasks (25 each adapted from SWE-bench Pro, SWE Atlas refactoring and DeepSWE) in which a tool-using simulated user reveals requirements progressively and inspects the agent's workspace; the original task verifiers, mean of two runs with Claude Opus 4.7 and GPT 5.5 as the user simulator; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 5 models tracked.

Top models

#ModelScore
1Claude Opus 4.8 (High)26.7
2GPT-5.5 (High)24.7
3Claude Sonnet 4.6 (High)18.8
4Gemini 3.5 Flash (High)17.3

Interactive version: theaggregate.ai/benchmark?slug=swe-interact · How It Works · Data refreshed daily, snapshot 2026-09-29.