Phind-CodeLlama-34B-v2: benchmark results
Provider: Other. Released 2023-08-28. Access: Open.
Unified ELO 1228 ± 14, rank #2602 of 2656 rated models, from 76 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| InfiBench | 59 | Score (%) | 94.3 |
| EvoEval Verbose | 72.56 | Pass@1 (%) | 91 |
| EvoEval Combine | 25 | Pass@1 (%) | 89 |
| EffiBench - NMU | 1.89 | Normalized Memory Usage | 87.8 |
| BigCode Models Leaderboard | 72 | HumanEval Python Pass@1 (%) | 84.7 |
| EvoEval Concise | 70.73 | Pass@1 (%) | 84 |
| MMLU-by-task - High School Statistics | 47.22 | Accuracy (%) | 83.4 |
| EvoEval | 52.13 | Pass@1 (%) | 74 |
| EvalPlus (HumanEval+ & MBPP+) | 67.1 | Pass@1 avg (%) | 73.4 |
| EvoEval Subtle | 63 | Pass@1 (%) | 72 |
| EvoEval Tool Use | 58 | Pass@1 (%) | 72 |
| EvoEval Creative | 35 | Pass@1 (%) | 70 |
Interactive version: theaggregate.ai/model?slug=phind-codellama-34b-v2 · How It Works · Data refreshed daily, snapshot 2026-09-19.