starchat2-15B-v0.1: benchmark results

Hugging Face H4's coding chat assistant, an SFT+DPO tune of StarCoder2-15B on UltraFeedback and Orca DPO pairs. Provider: HuggingFace. Released 2024-03-10. Access: Open.

Unified ELO 1472 ± 1, rank #829 of 1392 rated models, from 36 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HumanLikeness - Sound-270.12Humanlike Score (%)100
HumanLikeness - Syntax-186.93Humanlike Score (%)94.7
HumanLikeness - Word-224.52Humanlike Score (%)84.2
EvalPlus (HumanEval+ & MBPP+)67.9Pass@1 avg (%)76.6
HumanLikeness - Discourse-261.81Humanlike Score (%)68.4
HumanLikeness - Meaning-162.24Humanlike Score (%)68.4
RewardBench73.22Score (%)63.6
TuRTLe - Verilator Syntax87.81Average Score (%)61.4
HumanLikeness - Overall60.84Overall Humanlike (%)57.9
HumanLikeness - Syntax-273.32Humanlike Score (%)57.9
TuRTLe - Icarus Syntax86.51Average Score (%)56.8
HumanLikeness - Word-160.05Humanlike Score (%)52.6

Interactive version: theaggregate.ai/model?slug=starchat2-15b-v0-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.