Aider Refactoring Benchmark — leaderboard

Aider benchmark for model performance on code refactoring tasks.

Metric: Percent completed correctly (self-reported). Source: benchmarklist.com. Status: saturation imminent. 10 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet92.1
2O1 Preview75.3
3Claude 3 Opus72.3
4GPT-4o62.9
5GPT-450.6
6Gemini 1.5 Pro49.4
7O1 Mini44.9
8GPT-4 Turbo34.1
9DeepSeek V2.531.5

Interactive version: theaggregate.ai/benchmark?slug=aider-refactoring-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.