Qwen 2.5 Coder 32B Instruct — benchmark results
Alibaba's Apache-2.0 32B code specialist that matched GPT-4o on coding benchmarks - the open-source SOTA coder at release (November 2024). Provider: Alibaba. Released 2024-11-12. Access: Open.
Unified ELO 1506 ± 17, rank #778 of 1776 rated models, from 111 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open Japanese LLM - CG | 68.47 | Score (%) | 99.8 |
| EvalPlus (HumanEval+ & MBPP+) | 82.1 | Pass@1 avg (%) | 97.6 |
| BigCode Models Leaderboard | 83.2 | HumanEval Python Pass@1 (%) | 96.6 |
| Open LLM Leaderboard - MATH Level 5 | 49.55 | Score | 96.6 |
| GSM8K | 93 | Accuracy (%) | 95.7 |
| BigCodeBench | 49 | Pass@1 (%) | 95.2 |
| Open LLM Leaderboard - BBH | 52.27 | Score | 95 |
| EvalPlus | 82.1 | EvalPlus Avg. (self-reported) | 91.7 |
| MBPP+ | 77 | MBPP+ pass@1 (self-reported) | 91.7 |
| Open Japanese LLM - Janli Exact Match | 87.08 | Score (%) | 90.5 |
| HumanEval+ | 87.2 | HumanEval+ pass@1 (self-reported) | 89.1 |
| Open LLM Leaderboard - GPQA | 13.2 | Score | 87.4 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-coder-32b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.