Skywork-Reward-V2-Llama-3.1-8B: benchmark results

Provider: Skywork. Access: Open.

Unified ELO 1662 ± 27, rank #279 of 1629 rated models, from 23 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RewardBench 284.13Score (%)100
RewardBench 2 Focus98.38Score (%)100
RewardBench 2 Safety96.67Score (%)99.2
RewardBench 2 Factuality84.63Accuracy (%)97.8
Plan-RewardBench - Multi-Turn Planning (Easy)73.85Pairwise accuracy (%) on the 109 easy multi-turn planning pa95
RewardBench 2 Precise IF66.25Score (%)92.7
RewardBench 2 Ties81.24Score (%)92.3
RMGAP67.99Pairwise accuracy (%; chosen response scored above a rejecte87.5
RMGAP - Best-of-446.62Best-of-4 accuracy (%; chosen response scored above all thre87.5
RewardBench 2 Math77.6Score (%)84.3
Personalized RewardBench - Society and Culture71.88Pairwise accuracy (%) on the 1,074 Society and Culture test 78.1
Personalized RewardBench - Arts and Entertainment66.62Pairwise accuracy (%) on the 767 Arts and Entertainment test75

Interactive version: theaggregate.ai/model?slug=skywork-reward-v2-llama-3-1-8b · How It Works · Data refreshed daily, snapshot 2026-10-07.