WideSWE: leaderboard
Metric: Task success (%) over 120 cross-repository tasks (60 bug fixes, 60 features, 41 ecosystems); a task succeeds only when every target repository passes its hidden tests (Claude Code harness). Source: zju-aces-ise.github.io. Saturation forecast: Around 2032. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.8 Max (xHigh) | 37.5 |
| 2 | Claude Opus 5 (High) | 35 |
| 3 | GPT-5.6 Sol (High) | 32.5 |
| 4 | DeepSeek V4 Pro (High) | 26.67 |
Interactive version: theaggregate.ai/benchmark?slug=wideswe · How It Works · Data refreshed daily, snapshot 2026-10-01.