Llama-3-OffsetBias-RM-8B: benchmark results
Provider: Meta. Access: Open.
Unified ELO 1658 ± 45, rank #311 of 1605 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Focus | 95.96 | Score (%) | 96.2 |
| RewardBench | 89.42 | Score (%) | 91 |
| RewardBench Chat | 97.21 | Accuracy (%) | 86.9 |
| RewardBench Chat Hard | 81.8 | Accuracy (%) | 84.1 |
| RewardBench Reasoning | 91.92 | Accuracy (%) | 76.9 |
| RewardBench Precise IF | 40 | Score (%) | 71.9 |
| RewardBench Ties | 67.86 | Score (%) | 71.6 |
| RewardBench Safety | 86.76 | Accuracy (%) | 70.8 |
| Themis-CodeRewardBench | 71.76 | Preference accuracy (%; share of the benchmark's preference | 45 |
| Themis-CodeRewardBench - Functional Correctness | 79.8 | Preference accuracy (%; share of the benchmark's preference | 44.9 |
| Themis-CodeRewardBench - Security Hardness | 65.29 | Preference accuracy (%; share of the benchmark's preference | 44.9 |
| Themis-CodeRewardBench - Execution Efficiency | 59.05 | Preference accuracy (%; share of the benchmark's preference | 38.8 |
Interactive version: theaggregate.ai/model?slug=llama-3-offsetbias-rm-8b · How It Works · Data refreshed daily, snapshot 2026-09-26.