stablelm-zephyr-3B — benchmark results
Provider: Stability AI. Released 2023-11-21. Access: Open.
Unified ELO 1210 ± 73, rank #1675 of 1776 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench | 71.46 | Score (%) | 58 |
| Open LLM Leaderboard - MuSR | 9.79 | Score | 48.1 |
| RewardBench Chat Hard | 60.09 | Accuracy (%) | 46 |
| EuroEval Portuguese NLU - ScaLA PT | 8.98 | Linguistic acceptability Score (%) | 43.2 |
| EuroEval Portuguese NLU - MultiWikiQA PT | 65.92 | Reading comprehension Score (%) | 41.5 |
| EuroEval Portuguese NLU - SST-2 PT | 72.76 | Sentiment classification Score (%) | 40.7 |
| EuroEval Portuguese NLU | 45.29 | NLU Average Score (%) | 39.4 |
| RewardBench Reasoning | 75.73 | Accuracy (%) | 37.9 |
| Open LLM Leaderboard - IFEval | 36.83 | Score | 35.5 |
| EuroEval Portuguese Common Sense Reasoning | 14.18 | Common Sense Reasoning Average Score (%) | 35.3 |
| EuroEval Portuguese | 34.75 | Average Score (%) | 34.8 |
| RewardBench Safety | 74.05 | Accuracy (%) | 30.6 |
Interactive version: theaggregate.ai/model?slug=stablelm-zephyr-3b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.