starchat2-15B-v0.1 — benchmark results
Hugging Face H4's coding chat assistant, an SFT+DPO tune of StarCoder2-15B on UltraFeedback and Orca DPO pairs. Provider: HuggingFace. Released 2024-03-10. Access: Open.
Unified ELO 1422 ± 15, rank #1121 of 1776 rated models, from 36 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HumanLikeness - Sound-2 | 70.12 | Humanlike Score (%) | 100 |
| HumanLikeness - Syntax-1 | 86.93 | Humanlike Score (%) | 94.7 |
| HumanLikeness - Word-2 | 24.52 | Humanlike Score (%) | 84.2 |
| EvalPlus (HumanEval+ & MBPP+) | 67.9 | Pass@1 avg (%) | 76.6 |
| HumanLikeness - Discourse-2 | 61.81 | Humanlike Score (%) | 68.4 |
| HumanLikeness - Meaning-1 | 62.24 | Humanlike Score (%) | 68.4 |
| RewardBench | 73.22 | Score (%) | 63.6 |
| TuRTLe - Verilator Syntax | 87.81 | Average Score (%) | 60.5 |
| TuRTLe - Icarus Syntax | 86.51 | Average Score (%) | 58.1 |
| HumanLikeness - Overall | 60.84 | Overall Humanlike (%) | 57.9 |
| HumanLikeness - Syntax-2 | 73.32 | Humanlike Score (%) | 57.9 |
| HumanLikeness - Word-1 | 60.05 | Humanlike Score (%) | 52.6 |
Interactive version: theaggregate.ai/model?slug=starchat2-15b-v0-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.