InternLM2-7B-Reward: benchmark results
Provider: Shanghai AI Lab. Access: Open.
Unified ELO 1607 ± 49, rank #602 of 2656 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Chat | 99.16 | Accuracy (%) | 99.4 |
| RewardBench | 87.59 | Score (%) | 86.7 |
| RewardBench Reasoning | 94.53 | Accuracy (%) | 83.4 |
| RewardBench Safety | 87.16 | Accuracy (%) | 72.5 |
| RewardBench Precise IF | 40 | Score (%) | 71.9 |
| RewardBench Chat Hard | 69.52 | Accuracy (%) | 64.2 |
| RewardBench Ties | 51.64 | Score (%) | 43.2 |
| RewardBench Focus | 70.51 | Score (%) | 43.1 |
| RewardBench Math | 56.28 | Score (%) | 18.6 |
| RewardBench Factuality | 42.11 | Accuracy (%) | 8.2 |
| StoryAlign | 18.7 | Average (self-reported) | 7.1 |
Interactive version: theaggregate.ai/model?slug=internlm2-7b-reward · How It Works · Data refreshed daily, snapshot 2026-09-19.