Qwen 2.5 Coder 14B — benchmark results
Alibaba's Apache-2.0 14B code base model from the Qwen2.5-Coder series (November 2024), trained on 5.5T tokens of code-heavy data. Provider: Alibaba. Released 2024-11-12. Access: Open.
Unified ELO 1398 ± 46, rank #1239 of 1776 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GSM8K | 88.7 | Accuracy (%) | 89.9 |
| Open LLM Leaderboard - MMLU-Pro | 39.13 | Score | 84.1 |
| Open LLM Leaderboard - BBH | 40.52 | Score | 79.5 |
| Open LLM Leaderboard - MATH Level 5 | 22.51 | Score | 75.7 |
| MMLU | 75.2 | Accuracy (%) | 68.9 |
| HellaSwag | 80.2 | Accuracy (%) | 63.2 |
| ARC Challenge (AI2) | 66 | Accuracy (%) | 61.5 |
| WinoGrande | 76.8 | Accuracy (%) | 60 |
| Open Korean LLM Leaderboard | 114.14 | Average Score (%) | 51.9 |
| Open LLM Leaderboard - GPQA | 5.7 | Score | 47.8 |
| AI Energy Score (Text Generation) | 4 | Energy Score (1-5) | 47.5 |
| Open LLM Leaderboard - IFEval | 34.73 | Score | 33.2 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-coder-14b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.