Qwen 2.5 Coder 3B Instruct: benchmark results
Provider: Alibaba. Released 2024-11-12. Access: Open.
Unified ELO 1443 ± 23, rank #1020 of 1629 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Structured Local-Deployment MCQ | 75.67 | Strict accuracy (%): share of the 1,085 items answered with | 100 |
| GSM8K | 80.7 | Accuracy (%) | 77.4 |
| Delulu | 0.77 | pass@1 | 70 |
| C3-Bench (Controllable Code Completion) - Scale Control | 14.2 | Instruction-Following Rate (%; 909 SCC tasks: AST node-type | 65 |
| JSONSchemaBench | 91.2 | Schema Compliance (percentage) | 57.1 |
| Delulu - CodeBLEU | 47.2 | CodeBLEU (0-100): n-gram, syntax-tree and dataflow agreement | 55 |
| C3-Bench (Controllable Code Completion) - Implementation Control | 29.7 | Instruction-Following Rate (%; 1,286 ICC tasks: passes the u | 52.5 |
| Delulu - Edit Similarity | 69.3 | Edit similarity (0-100): character-level normalized Levensht | 50 |
| Delulu - Exact Match | 44 | Exact match (%) of the completion with the gold completion, | 50 |
| C3-Bench (Controllable Code Completion) - Implementation Control - Pass@1 | 40.5 | Pass@1 (%; unit tests on the 1,286 instructed ICC tasks) | 25 |
| CodeElo | 160 | Elo Rating | 23.5 |
| Aider Code Editing Leaderboard | 39.1 | Exercises completed correctly after one retry, pass_rate_2 ( | 21.5 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-coder-3b-instruct · How It Works · Data refreshed daily, snapshot 2026-10-07.