SmellBench Refactoring (Qwen Code) - Localization Accuracy: leaderboard

Metric: Localization accuracy (%): the edited class or method matches the ground-truth refactoring target; 294 injected-smell refactoring cases (7 smell types, 3 difficulty levels, guided and targeted instructions) in 7 real Python repositories, run once in the Harbor Docker framework with a 20-minute budget, Qwen Code agent harness; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.588.78
2GPT-5 Mini74.83
3Qwen 3 Coder 30B A3B Instruct67.01
4Qwen 3 Coder 480B A35B Instruct65.31
5Gemini 2.5 Flash63.95
6DeepSeek V3.249.66

Interactive version: theaggregate.ai/benchmark?slug=smellbench-refactoring-qwen-code-localization-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.