SWE Atlas - Refactoring: leaderboard

Scale SWE Atlas refactoring benchmark for coding agents that must restructure real code while preserving behavior and satisfying task-specific review criteria.

Metric: Score. Source: scale.com. Status: years away from saturation. 15 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)44.79
2GLM-5.242.38
3Gemini 3.1 Pro (Preview)33.81
4GLM-524.24
5Kimi K2.520.95
6MiniMax-M2.519.52
7Gemini 3 Flash10

Interactive version: theaggregate.ai/benchmark?slug=swe-atlas-refactoring · How It Works · Data refreshed daily, snapshot 2026-09-05.