Qwen 2.5 14B Instruct — benchmark results
Alibaba Qwen 2.5 14B instruction-tuned checkpoint. Provider: Alibaba. Released 2024-09-19. Access: Open.
Unified ELO 1518 ± 9, rank #734 of 1776 rated models, from 285 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - MATH Level 5 | 55.29 | Score | 99.1 |
| Open LLM Leaderboard - IFEval | 81.58 | Score | 98.7 |
| MT-Bench PL - Math | 8.1 | Judge Score (0-10) | 98 |
| Open Korean LLM Leaderboard | 768.68 | Average Score (%) | 97.6 |
| Open CoT - LSAT Analytical Reasoning | 11.74 | CoT Gain (%) | 96.6 |
| Finetuning with Scientific Data Increases Hall | 66.7 | OFS (self-reported) | 94.1 |
| Open Japanese LLM - Wiki Coreference SET F1 | 9.69 | Score (%) | 93.8 |
| OpenEval - BBQ | 92.59 | Exact Match (%) | 93.1 |
| OpenEval - IFEval Strict | 86.11 | Strict Accuracy (%) | 92.9 |
| SeaEval - Cultural Reasoning - SG-Eval (Zero-Shot) | 76.7 | Accuracy (%) | 92.9 |
| French LLM Leaderboard - IFEval FR | 66.42 | Score (%) | 92.6 |
| Open Arabic LLM - Aratrust Unfairness | 94.55 | Accuracy (%) | 91 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-14b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.