Themis-CodeRewardBench - Execution Efficiency: leaderboard

Metric: Preference accuracy (%; share of the benchmark's preference pairs in which the reward model scores the preferred code response above the rejected one, pointwise and reference-free; Execution Efficiency criterion, 1,309 pairs: runtime-improving commits, Pie4Perf, ECCO and EvalPerf). Source: arxiv.org. Saturation forecast: Estimated already saturated. 51 models tracked.

Top models

#ModelScore
1LDL-Reward-Gemma-2-27B-v0.165.47
2Llama-3.1-Tulu-3-70B-SFT-RM-RB265.16
3Skywork-Reward-V2-Qwen3-8B64.63
4Llama 3.1 70B Instruct RM RB264.02
5INF-ORM-Llama3.1-70B62.03
6ArmoRM-Llama3-8B-v0.161.96
7Llama 3.1 8B Base RM RB261.42
8URM-LLaMa-3.1-8B60.73
9InternLM2-20B-Reward60.2
10QRM-Llama3.1-8B-v259.74
11FsfairX-LLaMA3-RM-v0.159.74
12Llama-3-OffsetBias-RM-8B59.05
13GRM-llama3-8B-sftreg58.75
14InternLM2-7B-Reward57.37
15AceCodeRM-7B55.54

Interactive version: theaggregate.ai/benchmark?slug=themis-coderewardbench-execution-efficiency · How It Works · Data refreshed daily, snapshot 2026-09-26.