Qwen 2.5 Coder 32B — benchmark results
Alibaba Qwen 2.5 Coder 32B coding model row. Provider: Alibaba. Released 2024-11-12. Access: Open.
Unified ELO 1488 ± 35, rank #849 of 1776 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - MMLU-Pro | 47.81 | Score | 92.9 |
| GSM8K | 91.1 | Accuracy (%) | 92.6 |
| Open LLM Leaderboard - BBH | 48.51 | Score | 88.8 |
| Open LLM Leaderboard - GPQA | 12.86 | Score | 86.7 |
| Open LLM Leaderboard - MuSR | 15.87 | Score | 84.2 |
| Open LLM Leaderboard - MATH Level 5 | 30.89 | Score | 82.4 |
| FormalRewardBench | 37.6 | Pairwise Accuracy (self-reported) | 80 |
| MMLU | 79.1 | Accuracy (%) | 78.5 |
| GeoCode Leaderboard | 66.77 | AutoGEEval++ Pass@1 (self-reported) | 77.8 |
| HellaSwag | 83 | Accuracy (%) | 77.6 |
| WinoGrande | 80.8 | Accuracy (%) | 77.5 |
| BigCode Models Leaderboard | 57.1 | HumanEval Python Pass@1 (%) | 72.9 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-coder-32b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.