INF-ORM-Llama3.1-70B: benchmark results

Provider: Other. Access: Open.

Unified ELO 1851 ± 46, rank #54 of 1605 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RewardBench95.11Score (%)100
RewardBench Reasoning99.12Accuracy (%)100
RewardBench Chat Hard91.01Accuracy (%)99.4
RewardBench Safety96.44Accuracy (%)99.1
RewardBench Ties86.22Score (%)96.3
RewardBench Math69.95Score (%)88.3
RewardBench Focus90.3Score (%)86.7
RewardBench Factuality74.11Accuracy (%)83.7
RewardBench Precise IF41.88Score (%)83.2
RewardBench Chat96.65Accuracy (%)77.8
Themis-CodeRewardBench - Functional Correctness82.88Preference accuracy (%; share of the benchmark's preference 69.4
Themis-CodeRewardBench - Execution Efficiency62.03Preference accuracy (%; share of the benchmark's preference 65.3

Interactive version: theaggregate.ai/model?slug=inf-orm-llama3-1-70b · How It Works · Data refreshed daily, snapshot 2026-09-26.