Multi-SWE-Bench: leaderboard

Multilingual and multi-repository software issue benchmark evaluating whether LLMs can resolve bugs across codebases and languages.

Metric: Score (%). Source: llm-stats.com. Status: saturation imminent. 6 models tracked.

Top models

#ModelScore
1MiniMax-M2.752.7
2MiniMax-M2.551.3
3MiniMax-M2.149.4
4MiniMax-M236.2
5Qwen 3 Coder 480B A35B Instruct25.8

Interactive version: theaggregate.ai/benchmark?slug=multi-swe-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.