Themis-CodeRewardBench: leaderboard

Metric: Preference accuracy (%; share of the benchmark's preference pairs in which the reward model scores the preferred code response above the rejected one, pointwise and reference-free; all 8,866 pairs, pair-weighted over the five criteria). Source: arxiv.org. Saturation forecast: Estimated already saturated. 51 models tracked.

Top models

#ModelScore
1Skywork-Reward-V2-Qwen3-8B79.97
2Llama-3.1-Tulu-3-70B-SFT-RM-RB278.96
3Llama 3.1 70B Instruct RM RB278.23
4Llama 3.1 8B Base RM RB275.77
5LDL-Reward-Gemma-2-27B-v0.175.57
6INF-ORM-Llama3.1-70B74.84
7FsfairX-LLaMA3-RM-v0.173.52
8InternLM2-20B-Reward73.13
9QRM-Llama3.1-8B-v271.81
10GRM-llama3-8B-sftreg71.81
11ArmoRM-Llama3-8B-v0.171.76
12Llama-3-OffsetBias-RM-8B71.76
13InternLM2-7B-Reward71.73
14URM-LLaMa-3.1-8B71.62
15AceCodeRM-7B71.11

Interactive version: theaggregate.ai/benchmark?slug=themis-coderewardbench · How It Works · Data refreshed daily, snapshot 2026-09-26.