Llama-3.1-Tulu-3-70B-SFT-RM-RB2: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1730 ± 44, rank #178 of 1605 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RewardBench Factuality80.84Accuracy (%)96.4
RewardBench Ties83.08Score (%)93.7
RewardBench88.92Score (%)90.1
RewardBench Chat Hard82.68Accuracy (%)85.2
RewardBench Safety90.27Accuracy (%)85.2
RewardBench Chat96.93Accuracy (%)83
RewardBench Math67.76Score (%)82.1
Themis-CodeRewardBench - Readability and Maintainability73.58Preference accuracy (%; share of the benchmark's preference 82
Themis-CodeRewardBench - Execution Efficiency65.16Preference accuracy (%; share of the benchmark's preference 81.6
Themis-CodeRewardBench - Security Hardness73.76Preference accuracy (%; share of the benchmark's preference 81.6
Themis-CodeRewardBench - Functional Correctness86.23Preference accuracy (%; share of the benchmark's preference 80.6
Themis-CodeRewardBench78.96Preference accuracy (%; share of the benchmark's preference 80

Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-70b-sft-rm-rb2 · How It Works · Data refreshed daily, snapshot 2026-09-26.