Qwen2.5-Math-7B: benchmark results
Provider: Alibaba. Released 2024-09-16. Access: Open.
Unified ELO 1396 ± 1, rank #1152 of 1392 rated models, from 47 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PhysicsFinals | 100 | Score (self-reported) | 100 |
| Open LLM Leaderboard - MATH Level 5 | 30.51 | Score | 82.1 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 40.68 | Sentiment classification Score (%) | 54.1 |
| Open LLM Leaderboard - GPQA | 5.82 | Score | 48.7 |
| EuroEval Spanish NLU - ScaLA ES | 9.62 | Linguistic acceptability Score (%) | 45.3 |
| EuroEval Spanish NLU | 39.78 | NLU Average Score (%) | 42.4 |
| EuroEval Portuguese NLU - ScaLA PT | 7.68 | Linguistic acceptability Score (%) | 40.7 |
| EuroEval Spanish | 34.09 | Average Score (%) | 39.7 |
| EuroEval Portuguese | 37.62 | Average Score (%) | 39.2 |
| EuroEval Portuguese NLU - SST-2 PT | 72.07 | Sentiment classification Score (%) | 38.2 |
| EuroEval Norwegian Common Sense Reasoning | 29.52 | Common Sense Reasoning Average Score (%) | 38.1 |
| EuroEval Portuguese Common Sense Reasoning | 19.35 | Common Sense Reasoning Average Score (%) | 38.1 |
Interactive version: theaggregate.ai/model?slug=qwen2-5-math-7b · How It Works · Data refreshed daily, snapshot 2026-09-05.