Themis-CodeRewardBench - Memory Efficiency: leaderboard

Metric: Preference accuracy (%; share of the benchmark's preference pairs in which the reward model scores the preferred code response above the rejected one, pointwise and reference-free; Memory Efficiency criterion, 289 pairs: memory-improving commits and NoFunEval Memory). Source: arxiv.org. Saturation forecast: Estimated already saturated. 51 models tracked.

Top models

#ModelScore
1Skywork-Reward-V2-Qwen3-8B71.63
2Llama 3.1 8B Base RM RB269.9
3FsfairX-LLaMA3-RM-v0.167.82
4Llama-3.1-Tulu-3-70B-SFT-RM-RB267.13
5QRM-Llama3.1-8B-v266.78
6UltraRM-13B66.44
7InternLM2-20B-Reward65.74
8URM-LLaMa-3.1-8B65.05
9Llama 3.1 70B Instruct RM RB264.71
10Starling-RM-34B63.67
11INF-ORM-Llama3.1-70B62.63
12LDL-Reward-Gemma-2-27B-v0.162.63
13InternLM2-7B-Reward62.28
14internlm2-1.8B-reward62.28
15AceCodeRM-7B60.55

Interactive version: theaggregate.ai/benchmark?slug=themis-coderewardbench-memory-efficiency · How It Works · Data refreshed daily, snapshot 2026-09-26.