URM-LLaMa-3.1-8B: benchmark results
Provider: Other. Access: Open.
Unified ELO 1745 ± 40, rank #158 of 1605 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Focus | 97.58 | Score (%) | 99.5 |
| RewardBench | 92.94 | Score (%) | 96.3 |
| RewardBench Chat Hard | 88.16 | Accuracy (%) | 94.9 |
| RewardBench Precise IF | 45 | Score (%) | 93.1 |
| RewardBench Reasoning | 96.98 | Accuracy (%) | 92.9 |
| RewardBench Safety | 91.78 | Accuracy (%) | 91.2 |
| RewardBench Ties | 76.53 | Score (%) | 87.4 |
| RMGAP | 67.33 | Pairwise accuracy (%; chosen response scored above a rejecte | 75 |
| RMGAP - Best-of-4 | 45.28 | Best-of-4 accuracy (%; chosen response scored above all thre | 75 |
| RewardBench Math | 63.93 | Score (%) | 70.4 |
| PMDC | 0.07 | Bradley-Terry ranking score (log-strength; ArmoRM-Llama3-8B- | 66.7 |
| RewardBench Chat | 95.53 | Accuracy (%) | 63.9 |
Interactive version: theaggregate.ai/model?slug=urm-llama-3-1-8b · How It Works · Data refreshed daily, snapshot 2026-09-26.