Skywork-Reward-V2-Qwen3-8B: benchmark results

Provider: Skywork. Access: Open.

Unified ELO 1777 ± 44, rank #115 of 1605 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RMGAP69.97Pairwise accuracy (%; chosen response scored above a rejecte100
RMGAP - Best-of-449.16Best-of-4 accuracy (%; chosen response scored above all thre100
RewardBench Math77.05Score (%)98
RewardBench Focus96.36Score (%)97.4
RewardBench Safety94Accuracy (%)96.9
RewardBench Precise IF50Score (%)96.4
RewardBench Factuality79.89Accuracy (%)95.9
Themis-CodeRewardBench79.97Preference accuracy (%; share of the benchmark's preference 88
Themis-CodeRewardBench - Readability and Maintainability75.05Preference accuracy (%; share of the benchmark's preference 86
Themis-CodeRewardBench - Functional Correctness87.25Preference accuracy (%; share of the benchmark's preference 85.7
Themis-CodeRewardBench - Security Hardness74.97Preference accuracy (%; share of the benchmark's preference 85.7
Themis-CodeRewardBench - Memory Efficiency71.63Preference accuracy (%; share of the benchmark's preference 81.6

Interactive version: theaggregate.ai/model?slug=skywork-reward-v2-qwen3-8b · How It Works · Data refreshed daily, snapshot 2026-09-26.