Themis-CodeRewardBench - Functional Correctness: leaderboard
Metric: Preference accuracy (%; share of the benchmark's preference pairs in which the reward model scores the preferred code response above the rejected one, pointwise and reference-free; Functional Correctness criterion, 4,778 pairs: commit bug fixes, HumanEvalPack, MBPPPlusFix-Hard, MDEval, DebugEval and RunBugRun-V1). Source: arxiv.org. Saturation forecast: Estimated already saturated. 50 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Skywork-Reward-V2-Qwen3-8B | 87.25 |
| 2 | Llama 3.1 70B Instruct RM RB2 | 86.23 |
| 3 | Llama-3.1-Tulu-3-70B-SFT-RM-RB2 | 86.23 |
| 4 | LDL-Reward-Gemma-2-27B-v0.1 | 84.28 |
| 5 | INF-ORM-Llama3.1-70B | 82.88 |
| 6 | Llama 3.1 8B Base RM RB2 | 82.57 |
| 7 | AceCodeRM-7B | 82.48 |
| 8 | InternLM2-20B-Reward | 81.38 |
| 9 | FsfairX-LLaMA3-RM-v0.1 | 81.02 |
| 10 | QRM-Llama3.1-8B-v2 | 80.16 |
| 11 | URM-LLaMa-3.1-8B | 79.91 |
| 12 | Llama-3-OffsetBias-RM-8B | 79.8 |
| 13 | ArmoRM-Llama3-8B-v0.1 | 79.59 |
| 14 | InternLM2-7B-Reward | 79.3 |
| 15 | GRM-llama3-8B-sftreg | 79.05 |
Interactive version: theaggregate.ai/benchmark?slug=themis-coderewardbench-functional-correctness · How It Works · Data refreshed daily, snapshot 2026-09-26.