OpenCoder-8B-Instruct: benchmark results
Provider: Other. Released 2024-11-07. Access: Open.
Unified ELO 1367 ± 1, rank #1226 of 1392 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EvalPlus (HumanEval+ & MBPP+) | 74.4 | Pass@1 avg (%) | 88.7 |
| BigCodeBench | 43.2 | Pass@1 (%) | 69 |
| EvalPlus | 74.4 | EvalPlus Avg. (self-reported) | 50 |
| MBPP+ | 71.4 | MBPP+ pass@1 (self-reported) | 50 |
| EuroEval Portuguese NLU - ScaLA PT | 7.5 | Linguistic acceptability Score (%) | 40 |
| HumanEval+ | 77.4 | HumanEval+ pass@1 (self-reported) | 37 |
| EuroEval Italian NLU - ScaLA IT | 9.48 | Linguistic acceptability Score (%) | 35.6 |
| EuroEval Italian NLU - Sentipolc16 | 42.93 | Sentiment classification Score (%) | 35.3 |
| FullStackBench en | 43.63 | Score (self-reported) | 33.3 |
| EuroEval Portuguese NLU - SST-2 PT | 67.12 | Sentiment classification Score (%) | 32 |
| EuroEval Dutch NLU - DBRD | 69.4 | Sentiment classification Score (%) | 27.2 |
| EuroEval Norwegian Common Sense Reasoning | 14.03 | Common Sense Reasoning Average Score (%) | 27.1 |
Interactive version: theaggregate.ai/model?slug=opencoder-8b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.