Themis-CodeRewardBench - Functional Correctness: leaderboard

Metric: Preference accuracy (%; share of the benchmark's preference pairs in which the reward model scores the preferred code response above the rejected one, pointwise and reference-free; Functional Correctness criterion, 4,778 pairs: commit bug fixes, HumanEvalPack, MBPPPlusFix-Hard, MDEval, DebugEval and RunBugRun-V1). Source: arxiv.org. Saturation forecast: Estimated already saturated. 50 models tracked.

Top models

#ModelScore
1Skywork-Reward-V2-Qwen3-8B87.25
2Llama 3.1 70B Instruct RM RB286.23
3Llama-3.1-Tulu-3-70B-SFT-RM-RB286.23
4LDL-Reward-Gemma-2-27B-v0.184.28
5INF-ORM-Llama3.1-70B82.88
6Llama 3.1 8B Base RM RB282.57
7AceCodeRM-7B82.48
8InternLM2-20B-Reward81.38
9FsfairX-LLaMA3-RM-v0.181.02
10QRM-Llama3.1-8B-v280.16
11URM-LLaMa-3.1-8B79.91
12Llama-3-OffsetBias-RM-8B79.8
13ArmoRM-Llama3-8B-v0.179.59
14InternLM2-7B-Reward79.3
15GRM-llama3-8B-sftreg79.05

Interactive version: theaggregate.ai/benchmark?slug=themis-coderewardbench-functional-correctness · How It Works · Data refreshed daily, snapshot 2026-09-26.