Llama 3.1 8B Base RM RB2: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1647 ± 44, rank #331 of 1605 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RewardBench84.63Score (%)83.6
RewardBench Safety88.51Accuracy (%)79
RewardBench Chat Hard77.85Accuracy (%)77.8
Themis-CodeRewardBench - Memory Efficiency69.9Preference accuracy (%; share of the benchmark's preference 74.5
RewardBench Focus83.23Score (%)71.9
RewardBench Factuality72Accuracy (%)70.4
Themis-CodeRewardBench - Readability and Maintainability70.78Preference accuracy (%; share of the benchmark's preference 70
Themis-CodeRewardBench - Security Hardness71.24Preference accuracy (%; share of the benchmark's preference 69.4
Themis-CodeRewardBench75.77Preference accuracy (%; share of the benchmark's preference 68
Themis-CodeRewardBench - Functional Correctness82.57Preference accuracy (%; share of the benchmark's preference 65.3
Themis-CodeRewardBench - Execution Efficiency61.42Preference accuracy (%; share of the benchmark's preference 57.1
RewardBench Reasoning78.86Accuracy (%)48.5

Interactive version: theaggregate.ai/model?slug=llama-3-1-8b-base-rm-rb2 · How It Works · Data refreshed daily, snapshot 2026-09-26.