SmellBench Refactoring (Qwen Code) - Code Quality: leaderboard
Metric: Code quality score (0-100): readability and maintainability of the refactored code; an LLM judge (model not named) scores the refactoring against the smell analysis on a 0-10 rubric, normalized to 0-100; 294 injected-smell refactoring cases (7 smell types, 3 difficulty levels, guided and targeted instructions) in 7 real Python repositories, run once in the Harbor Docker framework with a 20-minute budget, Qwen Code agent harness; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 69.29 |
| 2 | GPT-5 Mini | 56.26 |
| 3 | Qwen 3 Coder 480B A35B Instruct | 45.73 |
| 4 | Qwen 3 Coder 30B A3B Instruct | 42.03 |
| 5 | Gemini 2.5 Flash | 41.63 |
| 6 | DeepSeek V3.2 | 36.77 |
Interactive version: theaggregate.ai/benchmark?slug=smellbench-refactoring-qwen-code-code-quality · How It Works · Data refreshed daily, snapshot 2026-09-29.