Themis-CodeRewardBench - Security Hardness: leaderboard

Metric: Preference accuracy (%; share of the benchmark's preference pairs in which the reward model scores the preferred code response above the rejected one, pointwise and reference-free; Security Hardness criterion, 991 pairs: vulnerability-fix commits, CodePrefBench Security, Vul4J, SecBench and NoFunEval Security). Source: arxiv.org. Saturation forecast: Estimated already saturated. 51 models tracked.

Top models

#ModelScore
1Skywork-Reward-V2-Qwen3-8B74.97
2Llama-3.1-Tulu-3-70B-SFT-RM-RB273.76
3Llama 3.1 70B Instruct RM RB272.96
4UltraRM-13B72.45
5Llama 3.1 8B Base RM RB271.24
6Starling-RM-34B68.72
7GRM-llama3-8B-sftreg67
8FsfairX-LLaMA3-RM-v0.166.6
9INF-ORM-Llama3.1-70B66.6
10InternLM2-20B-Reward66.26
11Llama-3-OffsetBias-RM-8B65.29
12InternLM2-7B-Reward63.57
13LDL-Reward-Gemma-2-27B-v0.163.17
14Eurus-RM-7B63.07
15QRM-Llama3.1-8B-v262.36

Interactive version: theaggregate.ai/benchmark?slug=themis-coderewardbench-security-hardness · How It Works · Data refreshed daily, snapshot 2026-09-26.