Skywork-Reward-Llama-3.1-8B-v0.2: benchmark results
Provider: Skywork. Access: Open.
Unified ELO 1808 ± 52, rank #142 of 2656 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Safety | 94.22 | Accuracy (%) | 97.5 |
| RewardBench | 93.13 | Score (%) | 96.9 |
| RewardBench Chat Hard | 88.38 | Accuracy (%) | 95.5 |
| RewardBench Focus | 94.14 | Score (%) | 93.6 |
| RewardBench Reasoning | 96.75 | Accuracy (%) | 91.7 |
| RewardBench Ties | 71.69 | Score (%) | 79.5 |
| RewardBench Precise IF | 40.62 | Score (%) | 77.3 |
| RewardBench Factuality | 69.68 | Accuracy (%) | 58.2 |
| RewardBench Chat | 94.69 | Accuracy (%) | 53.1 |
| StoryAlign | 31 | Average (self-reported) | 35.7 |
| RewardBench Math | 60.11 | Score (%) | 33.7 |
Interactive version: theaggregate.ai/model?slug=skywork-reward-llama-3-1-8b-v0-2 · How It Works · Data refreshed daily, snapshot 2026-09-19.