QRM-Gemma-2-27B: benchmark results

Provider: Other. Access: Open.

Unified ELO 1863 ± 42, rank #45 of 1605 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RewardBench94.44Score (%)99.4
RewardBench Safety95.78Accuracy (%)98.8
RewardBench Reasoning98.26Accuracy (%)98.2
RewardBench Chat Hard90.13Accuracy (%)98
RewardBench Focus95.35Score (%)94.9
RewardBench Factuality78.53Accuracy (%)94.4
RewardBench Ties83.21Score (%)94.2
PMDC1.27Bradley-Terry ranking score (log-strength; ArmoRM-Llama3-8B-88.9
RewardBench Math69.95Score (%)88.3
RewardBench Chat96.65Accuracy (%)77.8
RewardBench Precise IF37.19Score (%)49.5
Themis-CodeRewardBench - Execution Efficiency52.71Preference accuracy (%; share of the benchmark's preference 12.2

Interactive version: theaggregate.ai/model?slug=qrm-gemma-2-27b · How It Works · Data refreshed daily, snapshot 2026-09-26.