INF-ORM-Llama3.1-70B: benchmark results
Provider: Other. Access: Open.
Unified ELO 1851 ± 46, rank #54 of 1605 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench | 95.11 | Score (%) | 100 |
| RewardBench Reasoning | 99.12 | Accuracy (%) | 100 |
| RewardBench Chat Hard | 91.01 | Accuracy (%) | 99.4 |
| RewardBench Safety | 96.44 | Accuracy (%) | 99.1 |
| RewardBench Ties | 86.22 | Score (%) | 96.3 |
| RewardBench Math | 69.95 | Score (%) | 88.3 |
| RewardBench Focus | 90.3 | Score (%) | 86.7 |
| RewardBench Factuality | 74.11 | Accuracy (%) | 83.7 |
| RewardBench Precise IF | 41.88 | Score (%) | 83.2 |
| RewardBench Chat | 96.65 | Accuracy (%) | 77.8 |
| Themis-CodeRewardBench - Functional Correctness | 82.88 | Preference accuracy (%; share of the benchmark's preference | 69.4 |
| Themis-CodeRewardBench - Execution Efficiency | 62.03 | Preference accuracy (%; share of the benchmark's preference | 65.3 |
Interactive version: theaggregate.ai/model?slug=inf-orm-llama3-1-70b · How It Works · Data refreshed daily, snapshot 2026-09-26.