Qwen2.5-Coder-1.5B: benchmark results
Provider: Alibaba. Released 2024-11-12. Access: Open.
Unified ELO 1337 ± 1, rank #1290 of 1392 rated models, from 28 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GSM8K | 65.8 | Accuracy (%) | 67.7 |
| MERA - ruHumanEval | 4.82 | pass@1 (%) | 36.5 |
| MERA - SimpleAr | 94.9 | EM (%) | 36.1 |
| MERA - BPS | 87.6 | Accuracy (%) | 33.2 |
| ARC Challenge (AI2) | 45.2 | Accuracy (%) | 31.6 |
| MERA - ruCodeEval | 1.95 | pass@1 (%) | 31.6 |
| MERA - RWSD | 52.31 | Accuracy (%) | 30.4 |
| MERA - ruModAr | 45.62 | EM (%) | 28.2 |
| MERA - ruMultiAr | 23.14 | EM (%) | 27.5 |
| MMLU | 53.6 | Accuracy (%) | 23 |
| MERA - ruHateSpeech | 56.98 | Accuracy (%) | 22.4 |
| MERA - MathLogicQA | 34.47 | Accuracy (%) | 21.1 |
Interactive version: theaggregate.ai/model?slug=qwen2-5-coder-1-5b · How It Works · Data refreshed daily, snapshot 2026-09-05.