Skywork-Reward-V2-Qwen3-8B: benchmark results
Provider: Skywork. Access: Open.
Unified ELO 1777 ± 44, rank #115 of 1605 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RMGAP | 69.97 | Pairwise accuracy (%; chosen response scored above a rejecte | 100 |
| RMGAP - Best-of-4 | 49.16 | Best-of-4 accuracy (%; chosen response scored above all thre | 100 |
| RewardBench Math | 77.05 | Score (%) | 98 |
| RewardBench Focus | 96.36 | Score (%) | 97.4 |
| RewardBench Safety | 94 | Accuracy (%) | 96.9 |
| RewardBench Precise IF | 50 | Score (%) | 96.4 |
| RewardBench Factuality | 79.89 | Accuracy (%) | 95.9 |
| Themis-CodeRewardBench | 79.97 | Preference accuracy (%; share of the benchmark's preference | 88 |
| Themis-CodeRewardBench - Readability and Maintainability | 75.05 | Preference accuracy (%; share of the benchmark's preference | 86 |
| Themis-CodeRewardBench - Functional Correctness | 87.25 | Preference accuracy (%; share of the benchmark's preference | 85.7 |
| Themis-CodeRewardBench - Security Hardness | 74.97 | Preference accuracy (%; share of the benchmark's preference | 85.7 |
| Themis-CodeRewardBench - Memory Efficiency | 71.63 | Preference accuracy (%; share of the benchmark's preference | 81.6 |
Interactive version: theaggregate.ai/model?slug=skywork-reward-v2-qwen3-8b · How It Works · Data refreshed daily, snapshot 2026-09-26.