Qwen 2.5 Coder 7B Instruct: benchmark results
Alibaba's Apache-2.0 7B code specialist launched with Qwen2.5-Coder in September 2024, trained on 5.5T code-heavy tokens with 128K context. Provider: Alibaba. Released 2024-09-19. Access: Open.
Unified ELO 1449 ± 1, rank #948 of 1392 rated models, from 108 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GSM8K | 86.7 | Accuracy (%) | 87.6 |
| Open LLM Leaderboard - MATH Level 5 | 37.16 | Score | 87.5 |
| MERA - RCB | 58.45 | Accuracy (%) | 84.2 |
| MERA - BPS | 99.4 | Accuracy (%) | 84.1 |
| Open Korean LLM Leaderboard | 41 | Average Score (%) | 77.6 |
| Open LLM Leaderboard - IFEval | 61.47 | Score | 73.6 |
| DuckDB-NSQL | 57.3 | Execution Accuracy (%) | 66.1 |
| MERA - ruCodeEval | 20.67 | pass@1 (%) | 64.1 |
| MERA - ruHateSpeech | 78.11 | Accuracy (%) | 60.1 |
| BigCodeBench | 40.4 | Pass@1 (%) | 59.9 |
| MERA - ruHumanEval | 17.32 | pass@1 (%) | 58.7 |
| Open Japanese LLM - Mbpp Pylint Check | 49.2 | Score (%) | 57.8 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-coder-7b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.