Llama-3.1-Tulu-3-8B-RM — benchmark results
Provider: Meta. Released 2024-11-20. Access: Open.
Unified ELO 1367 ± 61, rank #1355 of 1776 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Factuality | 74.53 | Accuracy (%) | 86.5 |
| RewardBench Math | 64.48 | Score (%) | 73.7 |
| RewardBench Ties | 52.43 | Score (%) | 44.2 |
| RewardBench Safety | 74.22 | Accuracy (%) | 30.9 |
| RewardBench Precise IF | 34.69 | Score (%) | 29.6 |
| Open LLM Leaderboard - MuSR | 4.25 | Score | 20.9 |
| RewardBench | 59 | Score (%) | 20.7 |
| RewardBench Focus | 53.64 | Score (%) | 17.3 |
| Open LLM Leaderboard - GPQA | 0.89 | Score | 9 |
| Open LLM Leaderboard - IFEval | 16.7 | Score | 7.5 |
| Open LLM Leaderboard - BBH | 2.65 | Score | 4.3 |
| Open LLM Leaderboard - MATH Level 5 | 0 | Score | 1.3 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-tulu-3-8b-rm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.