Qwen 2.5 Coder 3B Instruct: benchmark results

Provider: Alibaba. Released 2024-11-12. Access: Open.

Unified ELO 1443 ± 23, rank #1020 of 1629 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Structured Local-Deployment MCQ75.67Strict accuracy (%): share of the 1,085 items answered with 100
GSM8K80.7Accuracy (%)77.4
Delulu0.77pass@170
C3-Bench (Controllable Code Completion) - Scale Control14.2Instruction-Following Rate (%; 909 SCC tasks: AST node-type 65
JSONSchemaBench91.2Schema Compliance (percentage)57.1
Delulu - CodeBLEU47.2CodeBLEU (0-100): n-gram, syntax-tree and dataflow agreement55
C3-Bench (Controllable Code Completion) - Implementation Control29.7Instruction-Following Rate (%; 1,286 ICC tasks: passes the u52.5
Delulu - Edit Similarity69.3Edit similarity (0-100): character-level normalized Levensht50
Delulu - Exact Match44Exact match (%) of the completion with the gold completion, 50
C3-Bench (Controllable Code Completion) - Implementation Control - Pass@140.5Pass@1 (%; unit tests on the 1,286 instructed ICC tasks)25
CodeElo160Elo Rating23.5
Aider Code Editing Leaderboard39.1Exercises completed correctly after one retry, pass_rate_2 (21.5

Interactive version: theaggregate.ai/model?slug=qwen-2-5-coder-3b-instruct · How It Works · Data refreshed daily, snapshot 2026-10-07.