Llama-3.1-Tulu-3-70B-SFT-RM-RB2: benchmark results
Provider: Meta. Access: Open.
Unified ELO 1730 ± 44, rank #178 of 1605 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Factuality | 80.84 | Accuracy (%) | 96.4 |
| RewardBench Ties | 83.08 | Score (%) | 93.7 |
| RewardBench | 88.92 | Score (%) | 90.1 |
| RewardBench Chat Hard | 82.68 | Accuracy (%) | 85.2 |
| RewardBench Safety | 90.27 | Accuracy (%) | 85.2 |
| RewardBench Chat | 96.93 | Accuracy (%) | 83 |
| RewardBench Math | 67.76 | Score (%) | 82.1 |
| Themis-CodeRewardBench - Readability and Maintainability | 73.58 | Preference accuracy (%; share of the benchmark's preference | 82 |
| Themis-CodeRewardBench - Execution Efficiency | 65.16 | Preference accuracy (%; share of the benchmark's preference | 81.6 |
| Themis-CodeRewardBench - Security Hardness | 73.76 | Preference accuracy (%; share of the benchmark's preference | 81.6 |
| Themis-CodeRewardBench - Functional Correctness | 86.23 | Preference accuracy (%; share of the benchmark's preference | 80.6 |
| Themis-CodeRewardBench | 78.96 | Preference accuracy (%; share of the benchmark's preference | 80 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-70b-sft-rm-rb2 · How It Works · Data refreshed daily, snapshot 2026-09-26.