Llama-3.1-Tulu-3-8B-RM — benchmark results

Provider: Meta. Released 2024-11-20. Access: Open.

Unified ELO 1367 ± 61, rank #1355 of 1776 rated models, from 13 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RewardBench Factuality74.53Accuracy (%)86.5
RewardBench Math64.48Score (%)73.7
RewardBench Ties52.43Score (%)44.2
RewardBench Safety74.22Accuracy (%)30.9
RewardBench Precise IF34.69Score (%)29.6
Open LLM Leaderboard - MuSR4.25Score20.9
RewardBench59Score (%)20.7
RewardBench Focus53.64Score (%)17.3
Open LLM Leaderboard - GPQA0.89Score9
Open LLM Leaderboard - IFEval16.7Score7.5
Open LLM Leaderboard - BBH2.65Score4.3
Open LLM Leaderboard - MATH Level 50Score1.3

Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-8b-rm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.