RewardBench Prior Sets (0.5 weight) — leaderboard

Metric: Score (%). Source: huggingface.co. 105 models tracked.

Top models

#ModelScore
1ArmoRM-Llama3-8B-v0.174.29
2GPT-4 Turbo73.63
3GPT-4o (2024-05-13)72.62
4GPT-4 Preview (0125)70.85
5Llama 3 70B Instruct70.35
6Claude 3 Sonnet (20240229)69.63
7Gemini 1.5 Flash (001)69.37
8c4ai-command-r-plus69.24
9Claude 3 Haiku (20240307)66.35
10GPT-3.5 Turbo (0125)65.48
11Faro-Yi-9B-DPO63.95
12Llama 3 8B Instruct60.82
13Nous-Hermes-2-Mistral-7B-DPO55.5
14starchat2-15B-v0.155.25
15zephyr-7B-alpha53.53

Interactive version: theaggregate.ai/benchmark?slug=rewardbench-prior-sets-0-5-weight · How the rankings work · Data refreshed daily, snapshot 2026-07-22.