CodeQwen1.5-7B Chat: benchmark results
Chat-tuned CodeQwen1.5 7B checkpoint, kept separate from the base CodeQwen1.5 7B row. Provider: Alibaba. Released 2024-04-16. Access: Open.
Unified ELO 1435 ± 19, rank #1855 of 2928 rated models, from 20 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BigCode Models Leaderboard | 87.2 | HumanEval Python Pass@1 (%) | 99.2 |
| RedCode | 74.75 | Score (%) | 92.9 |
| Big Code Memorization - HumanEval-ET pass@1 | 60.98 | HumanEval-ET pass@1 (%) | 88.2 |
| EvalPlus (HumanEval+ & MBPP+) | 73.8 | Pass@1 avg (%) | 87.9 |
| Big Code Memorization - HumanEval pass@1 | 69.51 | HumanEval pass@1 (%) | 82.4 |
| Big Code Memorization - HumanEval pass@50 | 63.68 | HumanEval pass@50 (%) | 82.4 |
| Big Code Memorization - HumanEval-ET pass@50 | 55.66 | HumanEval-ET pass@50 (%) | 82.4 |
| BigCodeBench | 39.6 | Pass@1 (%) | 56.4 |
| RepoQA | 62.8 | Score (self-reported) | 56.2 |
| EvalPlus | 73.85 | EvalPlus Avg. (self-reported) | 45.8 |
| Open LLM Leaderboard v1 - GSM8K | 27.9 | Accuracy (%) (5-shot) | 45.6 |
| HumanEval+ | 78.7 | HumanEval+ pass@1 (self-reported) | 43.5 |
Interactive version: theaggregate.ai/model?slug=codeqwen1-5-7b-chat · How It Works · Data refreshed daily, snapshot 2026-09-23.