LDL-Reward-Gemma-2-27B-v0.1: benchmark results
Provider: Other. Access: Open.
Unified ELO 1803 ± 45, rank #87 of 1605 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench | 94.99 | Score (%) | 99.7 |
| RewardBench Reasoning | 99.03 | Accuracy (%) | 99.4 |
| RewardBench Chat Hard | 90.79 | Accuracy (%) | 98.9 |
| RewardBench Safety | 93.78 | Accuracy (%) | 96.6 |
| RewardBench Factuality | 75.58 | Accuracy (%) | 90.1 |
| RewardBench Focus | 91.31 | Score (%) | 88.3 |
| RewardBench Ties | 76.33 | Score (%) | 86.3 |
| Themis-CodeRewardBench - Execution Efficiency | 65.47 | Preference accuracy (%; share of the benchmark's preference | 83.7 |
| RewardBench Math | 64.48 | Score (%) | 73.7 |
| RewardBench Chat | 96.37 | Accuracy (%) | 73.3 |
| Themis-CodeRewardBench - Functional Correctness | 84.28 | Preference accuracy (%; share of the benchmark's preference | 71.4 |
| Themis-CodeRewardBench | 75.57 | Preference accuracy (%; share of the benchmark's preference | 64 |
Interactive version: theaggregate.ai/model?slug=ldl-reward-gemma-2-27b-v0-1 · How It Works · Data refreshed daily, snapshot 2026-09-26.