QRM-Llama3.1-8B-v2: benchmark results
Provider: Other. Access: Open.
Unified ELO 1774 ± 51, rank #196 of 2656 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Safety | 94.67 | Accuracy (%) | 98.1 |
| RewardBench | 93.14 | Score (%) | 97.2 |
| RewardBench Reasoning | 96.77 | Accuracy (%) | 92.3 |
| RewardBench Chat Hard | 86.84 | Accuracy (%) | 92 |
| RewardBench Focus | 89.09 | Score (%) | 83.4 |
| RewardBench Ties | 72.34 | Score (%) | 80.5 |
| RewardBench Precise IF | 40.62 | Score (%) | 77.3 |
| RewardBench Chat | 96.37 | Accuracy (%) | 73.3 |
| RewardBench Factuality | 66.53 | Accuracy (%) | 45.9 |
| RewardBench Math | 61.2 | Score (%) | 43.9 |
| StoryAlign | 14.9 | Average (self-reported) | 0 |
Interactive version: theaggregate.ai/model?slug=qrm-llama3-1-8b-v2 · How It Works · Data refreshed daily, snapshot 2026-09-19.