Qwen2.5-Math-1.5B: benchmark results
Provider: Alibaba. Released 2024-09-16. Access: Open.
Unified ELO 1333 ± 1, rank #1301 of 1392 rated models, from 62 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MERA - ruMultiAr | 33.69 | EM (%) | 57.4 |
| MERA - ruHumanEval | 14.09 | pass@1 (%) | 54.3 |
| MERA - ruCodeEval | 7.99 | pass@1 (%) | 46.9 |
| MERA - SimpleAr | 96.9 | EM (%) | 44.7 |
| PhysicsFinals | 16.4 | Score (self-reported) | 43.8 |
| MERA - ruModAr | 50.48 | EM (%) | 42.6 |
| MERA - LCS | 11.4 | Accuracy (%) | 37.3 |
| EuroEval Icelandic Knowledge | 3.42 | Knowledge Average Score (%) | 31.6 |
| MERA - MathLogicQA | 36.66 | Accuracy (%) | 29.2 |
| EuroEval Italian NLU - ScaLA IT | 5.26 | Linguistic acceptability Score (%) | 29.1 |
| MERA - ruHHH | 55.62 | Accuracy (%) | 25.5 |
| EuroEval Spanish NLU - ScaLA ES | 0.48 | Linguistic acceptability Score (%) | 24.5 |
Interactive version: theaggregate.ai/model?slug=qwen2-5-math-1-5b · How It Works · Data refreshed daily, snapshot 2026-09-05.