Skywork-Reward-V2-Llama-3.1-8B: benchmark results
Provider: Skywork. Access: Open.
Unified ELO 1662 ± 27, rank #279 of 1629 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench 2 | 84.13 | Score (%) | 100 |
| RewardBench 2 Focus | 98.38 | Score (%) | 100 |
| RewardBench 2 Safety | 96.67 | Score (%) | 99.2 |
| RewardBench 2 Factuality | 84.63 | Accuracy (%) | 97.8 |
| Plan-RewardBench - Multi-Turn Planning (Easy) | 73.85 | Pairwise accuracy (%) on the 109 easy multi-turn planning pa | 95 |
| RewardBench 2 Precise IF | 66.25 | Score (%) | 92.7 |
| RewardBench 2 Ties | 81.24 | Score (%) | 92.3 |
| RMGAP | 67.99 | Pairwise accuracy (%; chosen response scored above a rejecte | 87.5 |
| RMGAP - Best-of-4 | 46.62 | Best-of-4 accuracy (%; chosen response scored above all thre | 87.5 |
| RewardBench 2 Math | 77.6 | Score (%) | 84.3 |
| Personalized RewardBench - Society and Culture | 71.88 | Pairwise accuracy (%) on the 1,074 Society and Culture test | 78.1 |
| Personalized RewardBench - Arts and Entertainment | 66.62 | Pairwise accuracy (%) on the 767 Arts and Entertainment test | 75 |
Interactive version: theaggregate.ai/model?slug=skywork-reward-v2-llama-3-1-8b · How It Works · Data refreshed daily, snapshot 2026-10-07.