Llama 3.1 70B Instruct RM RB2: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1768 ± 44, rank #125 of 1605 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RewardBench Ties88.35Score (%)97.9
RewardBench Factuality81.26Accuracy (%)96.9
RewardBench90.21Score (%)92.6
RewardBench Math69.95Score (%)88.3
RewardBench Safety90.95Accuracy (%)88.3
RewardBench Chat Hard83.55Accuracy (%)86.4
RewardBench Precise IF41.88Score (%)83.2
Themis-CodeRewardBench - Functional Correctness86.23Preference accuracy (%; share of the benchmark's preference 80.6
RewardBench Focus86.46Score (%)79.6
RewardBench Chat96.65Accuracy (%)77.8
Themis-CodeRewardBench78.23Preference accuracy (%; share of the benchmark's preference 76
Themis-CodeRewardBench - Readability and Maintainability71.25Preference accuracy (%; share of the benchmark's preference 74

Interactive version: theaggregate.ai/model?slug=llama-3-1-70b-instruct-rm-rb2 · How It Works · Data refreshed daily, snapshot 2026-09-26.