internlm2-1.8B-reward: benchmark results
Provider: Shanghai AI Lab. Access: Open.
Unified ELO 1438 ± 46, rank #985 of 1605 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench | 82.17 | Score (%) | 78.7 |
| RewardBench Reasoning | 87.24 | Accuracy (%) | 64.5 |
| RewardBench Chat Hard | 66.23 | Accuracy (%) | 60.8 |
| RewardBench Safety | 81.62 | Accuracy (%) | 52.2 |
| RewardBench Chat | 93.58 | Accuracy (%) | 45.5 |
| RewardBench Precise IF | 36.25 | Score (%) | 42.1 |
| Themis-CodeRewardBench - Readability and Maintainability | 64.71 | Preference accuracy (%; share of the benchmark's preference | 41 |
| Themis-CodeRewardBench - Memory Efficiency | 62.28 | Preference accuracy (%; share of the benchmark's preference | 37.8 |
| Themis-CodeRewardBench - Security Hardness | 62.36 | Preference accuracy (%; share of the benchmark's preference | 31.6 |
| RewardBench Focus | 59.6 | Score (%) | 26.5 |
| Themis-CodeRewardBench | 63.84 | Preference accuracy (%; share of the benchmark's preference | 18 |
| Themis-CodeRewardBench - Functional Correctness | 67.06 | Preference accuracy (%; share of the benchmark's preference | 14.3 |
Interactive version: theaggregate.ai/model?slug=internlm2-1-8b-reward · How It Works · Data refreshed daily, snapshot 2026-09-26.