InternLM2-20B-Reward: benchmark results
Provider: Shanghai AI Lab. Access: Open.
Unified ELO 1650 ± 49, rank #458 of 2656 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Chat | 98.88 | Accuracy (%) | 98.9 |
| RewardBench | 90.16 | Score (%) | 92.3 |
| RewardBench Reasoning | 95.76 | Accuracy (%) | 87.6 |
| RewardBench Safety | 89.46 | Accuracy (%) | 81.5 |
| RewardBench Chat Hard | 76.54 | Accuracy (%) | 74.7 |
| RewardBench Ties | 54.83 | Score (%) | 47.4 |
| RewardBench Focus | 72.53 | Score (%) | 47.2 |
| RewardBench Precise IF | 36.25 | Score (%) | 42.1 |
| RewardBench Math | 57.38 | Score (%) | 20.9 |
| RewardBench Factuality | 55.58 | Accuracy (%) | 17.9 |
| StoryAlign | 21.3 | Average (self-reported) | 14.3 |
Interactive version: theaggregate.ai/model?slug=internlm2-20b-reward · How It Works · Data refreshed daily, snapshot 2026-09-19.