FsfairX-LLaMA3-RM-v0.1: benchmark results

Provider: Other. Access: Open.

Unified ELO 1608 ± 43, rank #428 of 1605 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RewardBench Chat99.44Accuracy (%)100
RewardBench Prior Sets (0.5 weight)74.92Score (%)97.1
RewardBench Precise IF41.88Score (%)83.2
RewardBench83.38Score (%)79.6
RewardBench Safety86.76Accuracy (%)70.8
RewardBench Ties66.47Score (%)67.9
Themis-CodeRewardBench - Memory Efficiency67.82Preference accuracy (%; share of the benchmark's preference 67.3
RewardBench Reasoning86.44Accuracy (%)61.5
RewardBench Math62.84Score (%)60.7
RewardBench Chat Hard65.13Accuracy (%)58
Themis-CodeRewardBench73.52Preference accuracy (%; share of the benchmark's preference 58
Themis-CodeRewardBench - Functional Correctness81.02Preference accuracy (%; share of the benchmark's preference 55.1

Interactive version: theaggregate.ai/model?slug=fsfairx-llama3-rm-v0-1 · How It Works · Data refreshed daily, snapshot 2026-09-26.