Qwen2.5-Math-7B — benchmark results
Provider: Alibaba. Released 2024-09-16. Access: Open.
Unified ELO 1314 ± 38, rank #1504 of 1776 rated models, from 47 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - MATH Level 5 | 30.51 | Score | 82.1 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 40.68 | Sentiment classification Score (%) | 54.1 |
| Open LLM Leaderboard - GPQA | 5.82 | Score | 48.6 |
| EuroEval Spanish NLU - ScaLA ES | 9.62 | Linguistic acceptability Score (%) | 45.3 |
| EuroEval Spanish NLU | 39.78 | NLU Average Score (%) | 42.4 |
| EuroEval Portuguese NLU - ScaLA PT | 7.68 | Linguistic acceptability Score (%) | 40.7 |
| EuroEval Spanish | 34.09 | Average Score (%) | 39.7 |
| EuroEval Portuguese | 37.62 | Average Score (%) | 39.2 |
| EuroEval Portuguese NLU - SST-2 PT | 72.07 | Sentiment classification Score (%) | 38.2 |
| EuroEval Norwegian Common Sense Reasoning | 29.52 | Common Sense Reasoning Average Score (%) | 38.1 |
| EuroEval Portuguese Common Sense Reasoning | 19.35 | Common Sense Reasoning Average Score (%) | 38.1 |
| EuroEval Spanish NLU - CoNLL ES | 54.68 | Named entity recognition Score (%) | 37.6 |
Interactive version: theaggregate.ai/model?slug=qwen2-5-math-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.