UltraRM-13B: benchmark results
Provider: Other. Access: Open.
Unified ELO 1291 ± 37, rank #1428 of 1605 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench Prior Sets (0.5 weight) | 72.94 | Score (%) | 90.4 |
| RewardBench Chat | 96.37 | Accuracy (%) | 73.3 |
| Themis-CodeRewardBench - Readability and Maintainability | 71.18 | Preference accuracy (%; share of the benchmark's preference | 72 |
| Themis-CodeRewardBench - Security Hardness | 72.45 | Preference accuracy (%; share of the benchmark's preference | 71.4 |
| Themis-CodeRewardBench - Memory Efficiency | 66.44 | Preference accuracy (%; share of the benchmark's preference | 60.2 |
| RewardBench | 69.03 | Score (%) | 53.4 |
| Open LLM Leaderboard v1 - TruthfulQA MC2 | 47.91 | MC2 (%) (0-shot) | 39.2 |
| RewardBench Chat Hard | 55.48 | Accuracy (%) | 34.7 |
| RewardBench Focus | 60.81 | Score (%) | 29.8 |
| Themis-CodeRewardBench | 70.43 | Preference accuracy (%; share of the benchmark's preference | 28 |
| Themis-CodeRewardBench - Functional Correctness | 74.17 | Preference accuracy (%; share of the benchmark's preference | 22.4 |
| RewardBench Precise IF | 33.12 | Score (%) | 21.7 |
Interactive version: theaggregate.ai/model?slug=ultrarm-13b · How It Works · Data refreshed daily, snapshot 2026-09-26.