Themis-CodeRewardBench - Readability and Maintainability: leaderboard

Metric: Preference accuracy (%; share of the benchmark's preference pairs in which the reward model scores the preferred code response above the rejected one, pointwise and reference-free; Readability and Maintainability criterion, 1,499 pairs: code-style commits and NoFunEval Maintain). Source: arxiv.org. Saturation forecast: Estimated already saturated. 51 models tracked.

Top models

#ModelScore
1Skywork-Reward-V2-Qwen3-8B75.05
2Llama-3.1-Tulu-3-70B-SFT-RM-RB273.58
3Llama 3.1 70B Instruct RM RB271.25
4UltraRM-13B71.18
5Llama 3.1 8B Base RM RB270.78
6INF-ORM-Llama3.1-70B68.18
7InternLM2-7B-Reward67.38
8FsfairX-LLaMA3-RM-v0.167.31
9LDL-Reward-Gemma-2-27B-v0.167.31
10GRM-llama3-8B-sftreg66.11
11ArmoRM-Llama3-8B-v0.164.71
12internlm2-1.8B-reward64.71
13Starling-RM-34B64.58
14InternLM2-20B-Reward64.11
15Llama-3-OffsetBias-RM-8B63.91

Interactive version: theaggregate.ai/benchmark?slug=themis-coderewardbench-readability-and-maintainability · How It Works · Data refreshed daily, snapshot 2026-09-26.