SmellBench Refactoring (Qwen Code): leaderboard

Metric: Smell elimination score (0-100): how completely the injected code smell is removed; an LLM judge (model not named) scores the refactoring against the smell analysis on a 0-10 rubric, normalized to 0-100; 294 injected-smell refactoring cases (7 smell types, 3 difficulty levels, guided and targeted instructions) in 7 real Python repositories, run once in the Harbor Docker framework with a 20-minute budget, Qwen Code agent harness; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.550.34
2GPT-5 Mini32.55
3Qwen 3 Coder 480B A35B Instruct28.67
4DeepSeek V3.225.03
5Gemini 2.5 Flash22.79
6Qwen 3 Coder 30B A3B Instruct22.51

Interactive version: theaggregate.ai/benchmark?slug=smellbench-refactoring-qwen-code · How It Works · Data refreshed daily, snapshot 2026-09-29.